r/datascience 2d ago

Weekly Entering & Transitioning - Thread 20 Jul, 2026 - 27 Jul, 2026

8 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 1d ago

AI Structured Evaluation Pipelines to Improve Your AI Workflows

Thumbnail
heltweg.org
5 Upvotes

r/datascience 1d ago

Discussion How do you debug a forecasting model today when the error is quite bad?

0 Upvotes

This is for a personal study that will end up becoming an in-depth article and possibly a fully open source solution ideally without the AI slop that we see these days.

Let's say you’ve trained a model and the result is worse than the business wants. What do you check next?

Do you break the error down by customer, product, location, or individual series? Check if it gets worse at longer horizons? Look for bias, volatility, intermittent demand or outliers?

Go back to the backtesting setup, metric, or baseline? Or do you usually start trying other models?

Also do the tools you use make this easy or do you end up building custom notebooks, tables, and plots every time?

Thinking about the last time this happened:

  • What did you check first?
  • What actually helped you find the problem?
  • What did you have to build yourself?
  • Did you end up changing the model, data, validation setup, metric, or business expectation?

I’m trying to understand how people diagnose bad forecasts beyond comparing one overall error score against another.

EDIT/UPDATE because it seems like this is not clear enough:

I’m not looking for an if-else checklist that can explain why any forecast is bad. The answer obviously depends on the data, objective, validation setup and the decision the model is supposed to support.

I’m exploring if there is room for a small open-source tool around forecast evaluation. Before building anything, I’m trying to understand which checks people repeatedly run after they already have predictions, what they still build manually, and what existing tools already handle well.

So I’m mainly interested in specific workflows from projects rather than a general formula for fixing a model.


r/datascience 2d ago

Discussion Why Reddit Data Scientists Keep Saying Not To Use Prophet

Thumbnail
codebynight.dev
120 Upvotes

Couple thoughts and a small experiment to see why reddit hates prophet xD


r/datascience 2d ago

AI How to control reasoning effort and thinking-token budgets in LLMs

Thumbnail
magazine.sebastianraschka.com
6 Upvotes

r/datascience 4d ago

ML Inkling, a new open-weight 975B mixture-of-experts model, comes with a few surprises

Thumbnail
sebastianraschka.com
46 Upvotes

r/datascience 5d ago

Tools cosmos.gl, a WebGL library for visualizing network graphs

Thumbnail cosmos.gl
14 Upvotes

r/datascience 5d ago

AI Context degradation in LLMs: what the papers actually show, and the habits I built for long analysis sessions

Thumbnail
towardsdatascience.com
14 Upvotes

r/datascience 5d ago

Discussion How do sell Training Data?

Thumbnail
0 Upvotes

r/datascience 5d ago

Career | US Data scientist in pharma trying to figure out what’s the best path forward

50 Upvotes

I’m a data scientist (more like an analytics engineer) in pharma. My background is clinical - I went to school for a healthcare degree and then went into research before coming into data. With that, I have about a decade of experience now doing a little bit of a lot of things - statistics, epidemiology, data/analytics engineering, data visualisation, product management, data governance, etc.

Over the course of the last few years, I’ve felt a little stagnant - currently I’m not really doing anything that feels all that important lol, I mean essentially my team was building data pipelines and im redesigning the process for projects that are already ongoing or near completion. The only good thing is I have some downtime which gives me an opportunity to explore different teams and projects. There’s 3 teams I’d like to work with but I can’t work with them all at the same time and have to figure out how best to prioritise each team and allocate time so I can get better exposure

  1. Team 1 - a data science team that focuses on a specific disease area, the opportunity would be to continue working in data science while staying close to the business side of our projects by developing a deeper understanding of clinical context.

  2. Team 2 - Generative AI engineering - this would be more technical and I probably can’t work with this team right away until I get up to speed with learning concepts like embeddings, chunking, RAGs which I’ve never done before.

  3. Team 3 - the downstream users of my data pipelines who apply ML/AI techniques to the data for biomarker discovery

Just wanted to hear insights in terms of an industry perspective, which teams would be the best to work on a project with


r/datascience 6d ago

Challenges As a data scientist do you experiment with tools (open source or not) that solve specific issues around DS work? If yes, how do you think about uploading work data into those tools?

6 Upvotes

the context is that I am exploring a few recurring problems to solve especially around forecasting and working with time series data but setup a simple open source project around those.

my question is primarily about how is everyone handling their official datasets when trying new tools - do you not care, do you remove any identifiers then upload, do you create synthetic data with exactly same properties as the og dataset?

happy to answer more questions if this is not clear enough.


r/datascience 7d ago

Career | Europe I’m not ready

86 Upvotes

Three years ago I was able to pivot from engineering (no coding) to data science. I’ve been working here at a job I love, with an awesome team & boss and with a great pay. I’m 47, so not the youngest.

Now for family reasons I must leave and move back to my country of origin.

The thing is that, although I love the field and I keep reading books and trying to learn every day, this is such a vast field that I don’t think I’m ready at all. In these 3 years I’ve done basic ML projects, lots of xgboost, random forests, anomaly detection, dealing with pySpark with PB -sized dataframes, etc. But putting those models to production was made by my more experienced colleagues.

So, while I can say I’ve learnt on each project, I also see how MUCH I lack compared with my teammates, with 10+ years exp. on the field.

Today I just started looking for DS jobs and I just felt so depressed. Many ask for a DS who can do the whole thing from cleaning data to taking the models to production. I have no idea of that and it sounds extremely intimidating. I also dread the interview because while I can code, I often use LLMs for things that due to the pace of work, I simply decided to do with AI, i.e.: I understand and can read window functions but I’d need an AI to write them because I can never keep the sintaxis in mind. If an interviewer sees me struggling with the syntax, the interview is done

I feel very “green” to compete out there with other “proper” data scientists who have a well defined experience and knowledge. I wouldn’t mind applying to junior jobs but due to my age, most companies here wouldn’t hire a 47yr old “junior”.

I’m not able to work my old job in the location we’re moving to because it’s very niche and only a certain sector and a certain company size hires for that, and that isn’t there in our new area.

I don’t know what to do, or if this is normal and everybody feels like that. Or maybe you guys have any advice… Anything you can come up with will be highly appreciated because I need to provide for two kids and I just don’t know how. I’m beginning to feel desperate


r/datascience 7d ago

Discussion Value to the mentees?

0 Upvotes

Those received mentorship, did you find it worth your time in general?

In my early career, I actively participated in mentorship program at my alma mater. However, most if not all students wanted to work in tech. Considering I didn't (and still don't) work in tech, and that my employer was rarely hiring, I felt there was not much value I could provide. There's also this self-selection process at play, which is those who seek mentorship tends to already have a good idea of what they should be doing.

Now a decade into my career, my expertise is super irrelevant to people in a different industry. I'm also oblivious to the entry-level job market requirements.

Recently, my alma mater reached out again for mentors. I've skipped the last few requests but think maybe I should ask y'all before turning it down again.

To clarify, I'm not asking if it's worth it for me. I'm wondering if it's worth it for the students to speak with someone unfamiliar with entry-level job market, has expertise in a not niche but definitely not popular domain, and definitely don't lead to job opportunities.

Edit: I messed up somewhere that two posts were created. 0_o


r/datascience 7d ago

AI 5 trends that defined AI engineering at World's Fair 2026

Thumbnail
latent.space
2 Upvotes

r/datascience 7d ago

Discussion This psychology study on why some people are more impressed by corporate buzzwords has nothing to do with AI, yet it immediately reminded me of what I've been seeing in data science since ChatGPT took off

91 Upvotes

I know buzzwords have always been common in business, but AI has taken it to another level. AI is being talked about everywhere, and consultants are pitching executives with outrageous claims about how it will revolutionize everything.

I came across an interesting article today, and one quote really caught my attention:

https://www.psypost.org/new-study-finds-link-between-receptivity-to-corporate-bullshit-and-weaker-leadership-skills/

"Across the studies, Littrell found that individuals differed significantly in how impressed they were by corporate buzzword statements. Those with higher corporate-bullshit receptivity scores were more likely to view jargon-heavy statements as insightful or indicative of business expertise. They were also more likely to engage in persuasive 'bullshitting' themselves, using exaggerated or misleading language to impress others.

At the same time, higher receptivity was associated with lower scores on measures of analytic thinking and fluid intelligence, suggesting that individuals who were more impressed by corporate jargon were also less likely to critically evaluate information."

It made me wonder if we're seeing this play out with AI and data science.

The AI boom has created an absolute paradise for people who are great at talking about tech, but don't actually build.

For those of you working in data science, analytics, or ML, have you noticed this in your company or with clients? Has the GenAI hype changed how technical decisions get made, or is this just the same corporate behavior we've always had with a new set of buzzwords?


r/datascience 8d ago

Tools A guide to profiling attention layers in PyTorch, part 3 of a series

Thumbnail
huggingface.co
12 Upvotes

r/datascience 9d ago

Weekly Entering & Transitioning - Thread 13 Jul, 2026 - 20 Jul, 2026

10 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 10d ago

Discussion Snowflake Python question about StandardScaler function

9 Upvotes

I'm running the following code in Snowflake Python to standardize my training, evaluation, and test data prior to predictive modeling:

from snowflake.ml.modeling.preprocessing import StandardScaler

all_cols = df_train3.columns

target_col = "AB_POST"

passthrough_cols = ["SANHO", "SCNHO"]

scaler = StandardScaler(

input_cols=[c for c in all_cols if c not in [target_col] + passthrough_cols],

output_cols=[c for c in all_cols if c not in [target_col] + passthrough_cols], # Overwrite or create new

drop_input_cols=False # Set True to remove original unscaled columns

)

scaler.fit(df_train3)

train_df_scaled = scaler.transform(df_train3)

val_df_scaled = scaler.transform(df_eval3)

test_df_scaled = scaler.transform(df_test3)

I'm getting the following error when I run the code -- I'm not sure what this means:

Exception: Provided column names ['TOTAL_MH_CLASSES', 'STFLAG',..., 'ADS_FA_RISK_NEW'] does not index into the dataset.


r/datascience 11d ago

Statistics ARIMA Is Boring, and That Is Why I Still Like It

Thumbnail
codebynight.dev
172 Upvotes

r/datascience 12d ago

Analysis GPT 5.6 has 72 possible configurations. What's a good default?

Thumbnail
sebastianraschka.com
10 Upvotes

r/datascience 12d ago

Career | US All these layoffs have made me question my job search

204 Upvotes

I've been job hunting for a few months now, applying to big tech and startups. But seeing the recent Microsoft layoffs made me stop and ask myself what I'm actually looking for in a new job. Instability and more money?

Right now I'm at a company that hasn't done layoffs since maybe the financial crisis. I know how fortunate that is. But if I switch jobs, I could make an extra $50K. So I keep asking myself: is that extra 50K worth the instability that comes with tech jobs right now? What if I join a company and get laid off within a year?

What does everyone think of these layoffs? Despite record profits, there doesn't seem to be an end to them.


r/datascience 13d ago

Education Toto-2.0: Time Series Multivariate Forecasting Finally Scales Like LLMs

Thumbnail
aihorizonforecast.substack.com
51 Upvotes

r/datascience 14d ago

Discussion Skill engineering and the case against one-shot AI design

Thumbnail
latent.space
15 Upvotes

r/datascience 15d ago

Discussion Managing/ Dealing with Junior Data Scientists?

236 Upvotes

I've been in the 'data science' space for a decade+ or so now. One thing I've noticed is that generally - give or take - outside of the elite jobs (<2-3% aka not me and almost certainly not you) the caliber of coworkers has declined drastically.

I'm not some fabled data scientist. I wasn't some GitHub nerd who had everything embroil or terminal wizard nor could I write out the math to a GBM on a blackboard. I'd even forget basic obvious statistics.

But I felt like I had common sense.

Now I'm a manager/director. I work with data scientists. And I'm just generally freaked out by the absolute lack of basic common sense. This is across the last 7 that I have managed.

Examples include:

  1. Not visualizing or plotting the KPI/Target (sales). Not realizing there were no recorded sales on major holidays.
  2. Telling me everything is improving from a sales perspective that it's up 4%...... from period 1 vs period 2... when ignoring that period 2 had 6% more days so in fact it's worse.
  3. obscure models that are overkill and a bunch of statistics ive never heard of instead of just telling me that the impact of our promotions is declining.
  4. General sense of not knowing what is even rational (e.g., our marketing ROI $1023 - no its not lol)

As I begin to delegate more I begin to get more freaked out by what I see. I can't be presenting to clients such obvious insane mistakes. But these are the candidates and profiles that get forced upon me or the team I inherit.

Are there any best strategies for dealing with this? I want to be seen as someone who can 'develop' the team... not just saying people are useless, but such glaring mistakes are insane.

Yes, alot of these things are perhaps due to them being crunched for time, or not knowing what objective is, or being focused on other things. I'm not talking about those examples. I'm talking about like year 1-2 not day 1 employees, not doing basic data checks.

As a data scientist I was obsessed with finding bits of info or making sure things were right. Now it seem every common for people to copy and paste code into chatgpt and have no idea about anything else around it?


r/datascience 16d ago

Education Build a reasoning model from scratch, the new book is out

Thumbnail
sebastianraschka.com
33 Upvotes