The Tools I Used, 2025
A working data toolkit as it stood in early 2025, preserved as a dated snapshot with a 2026 note on what the list got wrong.
Tools don’t make the data professional. They do make us look like wizards, which is close enough most days. This was the working set in early 2025: the apps and IDEs that kept my code functional and my coffee intake just shy of lethal.
The ones I wouldn’t give up
SQL Server Management Studio. The grumpy old friend that judges your
SELECT * habits and still has your back. Perfect for “I just need to fix one
stored procedure” moments that turn into 3 AM.
DBeaver. Connects to everything: Databricks, Snowflake, MySQL, BigQuery, your neighbor’s spreadsheet. Free, which is its own argument.
VS Code. Small footprint, infinite plugins. Python, JSON, Markdown, and the persistent fiction that I’ll learn Rust. Extensions worth having: SQL Formatter, a Python linter, Error Lens.
Git. My code’s time machine. Lets me rewind to the era before I optimized the ETL pipeline into oblivion.
The workhorses
Apache Airflow. The cron job you wish you had. Also an invitation to name your DAGs after Tolkien characters, which you will regret at handover.
dbt. Turns SELECT statements into something testable. Writing Jinja is cheaper than therapy.
Airbyte. Duct tape for integration, for that one niche SaaS tool your CFO loves. Open source, so you stop paying enterprise rates for connectors that break on Tuesdays.
Docker. “Works on my machine” stops being a lie.
Postman. The API whisperer, and a quiet judge of undocumented endpoints.
Power BI and Tableau. Because stakeholders need a picture before they’ll accept why the data takes so long.
Analysis and modeling
pandas. The Excel replacement your laptop’s fan is praying you’ll stop using.
Polars. pandas’ faster cousin. Worth the switch the first time your CSV outgrows your RAM, and mildly depressing once you realize how much time you spent waiting on group-bys.
scikit-learn. Assemble a model with instructions a PM can almost follow.
TensorFlow and PyTorch. For when you want to train a network or cosplay as a PhD student. Debugging gradients is its own genre of suffering.
Jupyter. Where we write 90% prose, 9% code, and 1% #TODO: fix this later.
MLflow. Adult supervision for experiments, so nothing ends up named
v23_final_FINAL.ipynb.
XGBoost. If your Kaggle score is bad, this is the redemption arc.
The specialists
dbatools. PowerShell module that automates backups through migrations. For when right-clicking feels too manual.
Great Expectations. Data quality’s fun police. Catches the bad batch before it ruins a morning.
Obsidian. Where half-baked ideas go, filed under the generous heading of knowledge management.
GitHub Copilot. Writes the boilerplate. Ten out of ten for productivity, two out of ten for the existential bit.
Things I’d tell myself
Don’t over-tool. A hammer is great until you use it on a lightbulb.
If your model hits 99.9% accuracy, you’re either a genius or you forgot to split your data. Check for leakage.
That #TODO: handle NaN comment will outlive you. Write the code.
Version your notebooks. Losing eight hours to a kernel crash should be illegal.
A note from 2026
I’m keeping this one exactly as written, because the gap between it and now is more interesting than the list was.
Look at the AI section. One entry: Copilot, filed under “pro tips,” described as a boilerplate generator. That was a reasonable read at the time and it was wrong within a year. The thing I now use for hours a day, in a terminal, against real repositories, doesn’t appear on this list because the category barely existed when I wrote it.
A few other absences that date it precisely. Nothing on open table formats, which turned out to matter more than most of the tools sitting on top of them. Nothing on retrieval or evaluation, because I hadn’t built anything that needed them yet. And no catalog or governance tooling at all, which is the gap that ended up mattering most.
The durable part is unglamorous and sits at the top: SSMS, DBeaver, VS Code, Git, Docker. Five years of churn underneath and the things you actually touch every day barely moved.
I’m not updating this list in place. The current one lives separately, and keeping both is the point: in another two years that one will read the same way this one does now.
This is older work.
Current writing lives in the main feed, where the thinking has moved on from most of what is here.