Skip to content

Repository files navigation

✦⊹ Data Center Drive Reliability Analytics and Replacement Intelligence System ✦⊹

Turning 10 million+ hard drive health records into actionable business insights.


✦ About this Project

Hard drives fail... but what if we could spot the warning signs before they fail?

In this project, I analyzed real production hard drive data from Backblaze using SMART health attributes to study drive reliability, compare failure rates across manufacturers and datacenters, identify risky drives, estimate business impact, and build a replacement priority list.

Rather than reacting after failures happen, this project focuses on predictive maintenance and smarter storage management through data analytics and interactive Power BI dashboards..


✦ Dataset

This project uses the publicly available Backblaze Hard Drive SMART Dataset collected from production data centers.

Official Dataset Page

https://www.backblaze.com/cloud-storage/resources/hard-drive-test-data

Dataset Used: Backblaze Drive Stats – Q4 2025 (December 2025 daily CSV files)

https://f001.backblazeb2.com/file/Backblaze-Hard-Drive-Data/data_Q4_2025.zip

Only the December 2025 daily files were analyzed in this project.


✦ Dataset Summary

✦ 31 Daily CSV Files

✦ 10,488,607 Drive Day Records

✦ 340,501 Unique Drives

✦ 78 Drive Models

✦ 6 Data Centers

✦ Around 3.84 GB Dataset Size


✦ What This Project Does

✦ Processes millions of real world SMART drive records

✦ Analyzes drive reliability across manufacturers and datacenters

✦ Detects high risk drives using SMART health indicators

✦ Generates an intelligent replacement priority list

✦ Estimates business cost and failure impact

✦ Visualizes insights through an interactive Power BI dashboard


✦ Dashboard Highlights

⭐ Executive Overview Dashboard

drive_reliability_dashboard

✦ Technologies Used

Python • Pandas • NumPy • Google Colab • Power BI


✦ Project Workflow

Dataset → Data Cleaning → Reliability Analysis → Risk Detection → Business Insights → Power BI Dashboard


✦ Key Results

✓ Processed 10.4M+ drive day-to-day records

✓ Analyzed 340K+ production hard drives

✓ Identified 152 drive failures

✓ 22K+ risky live drives flagged using SMART indicators

✓ Built an intelligent replacement priority list

✓ Estimated business cost exposure

✓ Designed 3 interactive Power BI dashboards


✦ Why This Project Matters

Storage failures can lead to downtime, data loss, and expensive replacements.

This project demonstrates how real SMART telemetry can be transformed into meaningful analytics that help organizations monitor drive health, reduce failures, and make smarter maintenance decisions.


⊹ Author

Manogna

⭑✮ If you enjoyed this project or have suggestions, feel free to connect. I'm always happy to learn and collaborate ₊˚⊹

About

Analyzed 10M+ real hard drive records to find why drives fail, spot risky drives early, decide which drives need attention first, estimate possible costs, and show the results in Power BI dashboards.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages