AbruvaTech

AbruvaTech

Share

I was 2 years into a PhD in game theory. Then I became homeless. 18 months later, I was a Senior Data Scientist at Mastercard. Mathematician.

Former Ministry of Finance. Lead Instructor at BrainStation. I teach the system I used to rebuild
www.abruva.com

09/10/2026

Some of you won't like this... but tutorials are wasting your time. ⏳
Watching β‰  learning. Doing = learning.
πŸ“ Learning SQL? Learn it with a real project, solving an actual problem β€” comparing home prices across cities, optimizing suggested orders in an app.
πŸ“ Learning Python? Apply it to those exact same questions β€” then go further, like predicting house prices.
βœ… Learning GitHub? Do it while you build all of the above, not separately. Create a repo for every project. Push any code, even if it's one line.
⚠️ Tutorials feel like progress. A shipped project actually is progress.
That's the difference between people who talk about data science and people who get hired for it.
Save this and follow for more real ways to break into data science. 🍁

09/07/2026

The one thing that is killing your job search as a data scientist.

Last night I was at an LP concert. Her voice, her performance, the whole show β€” genuinely one of the best I've ever seen. And the first thing I thought was: why are her tickets so cheap? Why doesn't everyone know about her?

Meanwhile, singers with way less talent are selling out arenas.

It's not talent. It's visibility.

πŸ“ LP is a full package. Unique voice, real artistry, unforgettable live performance.
βœ… What she's missing isn't skill. It's the visibility other artists are actively building.
⚠️ That's the exact same reason your job search is stuck β€” you're not less qualified, you're less seen.

Being good at your job isn't enough if no one knows you exist. Visibility isn't optional in data science β€” it's part of the work.

Save this and follow for the data science lessons nobody else is sharing. 🍁

09/03/2026

Your model can fail and look exactly like success.

Target leakage isn't a bug you'll catch in testing β€” it's a feature that already knows the answer.

πŸ“ It happens when a feature is a duplicate (or near-duplicate) of what you're predicting
βœ… Ask: could this feature only exist after the outcome already happened
⚠️ If yes β€” it doesn't belong in your model, no matter how good your score looks

A near-perfect training score isn't a win. Sometimes it's the first warning sign.

Save this and follow for more data science made simple. 🍁 datascience Canada πŸ‡¨πŸ‡¦ USA πŸ‡ΊπŸ‡Έ Europe

08/29/2026

AI is making data scientists worse. Here's how to not be one of them.

AI is a tool β€” not a replacement for your judgment.

πŸ“ Use it to find recent research, summarize your own ideas for review, and speed up your code (especially plots)
⚠️ Never let it build your model, interpret your results, or post outputs you haven't reviewed
⚠️ Skip it for your emails β€” grammar checks are fine, but your thinking should stay yours

The data scientists who get ahead aren't the ones who use AI the most. They're the ones who still think for themselves.

Save this and follow for more data science concepts made simple. 🍁

08/27/2026

You taught your model to read tomorrow's newspaper. And you probably didn't even know you did it.
Most people split their data randomly β€” which works fine for static datasets. But if your data has a timeline, random splitting is a trap.
Here's what happens: you train your model on December fraud patterns, then test it on January data. The model isn't learning to predict the future. It's memorizing patterns that haven't happened yet in your test set.
The result? A model that scores 95% in your notebook and falls apart the moment it meets real, forward-moving data.
Fraud patterns evolve. New tactics show up every month. A model trained on next month's fraud will look brilliant on paper β€” and fail silently in production.
The fix is simple: Split by time, not at random. Train on the past. Test on what happens after. That's the only way your test score actually means something.
Random split = future leaking into the past.
If your data has a timeline and you're splitting randomly, you're not testing your model. You're testing how well it memorized the future.
Follow so you don't miss the rest of the series.

08/22/2026

Everyone gets this wrong. Training a model isn't coding. It's teaching.
Training a model is like teaching a new employee to spot fraud.
You hand them thousands of past transactions. Already labeled. Fraud. Not fraud. You already know the answers.
The employee studies both piles side by side. What do the fraud cases have in common? What do the normal ones have in common?
Same time of day. Same type of purchase. Same odd pattern. They start to see what actually separates fraud from normal.
That's what the model is really learning. Not the transactions themselves. The pattern that tells them apart.
Once it knows that pattern, you can hand it a brand new transaction it's never seen before β€” and it'll know which side of the line it belongs on.
That's a trained model.
Save this, and follow for more data science made simple.

UBCI

08/19/2026

I heard "inference" on the job for the first time and had no idea which kind they meant.

Turns out there are two.

Statistical inference: you have a sample, and you use it to make a claim about a much bigger population. Think surveying 500 shoppers to estimate what all shoppers likely do.

Model inference: your trained model making a prediction on new data. Think your fraud model scoring one new transaction β€” fraud, or not.

Same word. One is about understanding a population. The other is about scoring a single case.

Next time someone says "inference," ask which one they mean.

Save this for the next data science conversation that leaves you confused.

08/15/2026

UBCI UBC Sauder School of Business

08/15/2026

08/12/2026

Not every problem needs a predictive model.
Sometimes you just need to describe what happened β€” and that's where RFM comes in.
RFM is a descriptive model. It doesn't predict the future. It analyzes past customer behavior and tells you exactly who matters right now.
Here's how it works:
πŸ“ Recency (R): How recently did they buy?
πŸ“ Frequency (F): How often do they buy?
πŸ“ Monetary (M): How much do they spend?
Combine these three scores, and you can segment your customers:
βœ… High RFM = Your VIPs
⚠️ Low RFM = At risk of leaving
It won't tell you what customers will do next β€” but it will tell you who to focus on right now.
That's the power of descriptive analytics.
Save this and follow for more data science concepts made simple. 🍁

UBCI UBC Sauder School of Business

Want your business to be the top-listed Beauty Salon in Vancouver?
Click here to claim your Sponsored Listing.

Address


Vancouver, BC