■ HIGH RISK ■ Technology
AI is automating data science from both ends — AutoML builds the models, LLMs write the queries and the cleaning code — which guts the job's routine middle. The data scientists who survive are the ones who were really doing business problem-framing all along; the notebook jockeys are in trouble.
“Auto-ML builds models while you're still cleaning the dataset.”
Our AI replacement risk score — how we score jobs
The unglamorous truth of data science was always that most of the day is not modeling. It's meeting with a product manager to figure out what question is actually being asked, hunting for where the data lives, writing SQL, cleaning the mess, building features, and then — the short fun part — fitting models, followed by the long haul of dashboards, A/B test analysis, and explaining to stakeholders why the number went down. The XKCD-famous 80% data-janitorial share was the job's moat: tedious, but it required a human.
That moat is draining. LLM coding assistants now write competent SQL, pandas, and cleaning pipelines from a plain-English description; AutoML platforms handle feature selection, model search, and tuning; and 'chat with your data' BI tools let executives self-serve the simple questions that used to become Jira tickets. Ironically, data scientists built the machinery of their own exposure. The role is also squeezed organizationally: analytics engineering absorbed the pipeline work, ML engineering absorbed deployment, and LLM APIs replaced many bespoke NLP and classification models that once justified a modeling team.
What resists is everything around the code. Deciding what to measure, spotting that the dataset's real problem is selection bias rather than model choice, designing a valid experiment, telling a VP their pet metric is confounded — these demand statistical judgment, domain context, and organizational courage that autocomplete doesn't have. Causal inference, experimentation at scale, and ML product ownership are growing even as commodity modeling shrinks. The title likely survives; the junior version of it — the person hired to clean data and fit sklearn models — is the layer being automated away, which is exactly why our risk score sits at the anxious middle of the scale.
Automatability: our editorial assessment of current and near-term AI capability
The squeeze is already visible in junior hiring and intensifies through the late 2020s as coding assistants and AutoML become default tooling. By ~2030 expect the routine-analytics tier of the role to be largely absorbed by software and self-serve BI, while demand holds for experimentation, causal inference, and ML product roles. The occupation transforms sharply this decade; entry-level generalist positions feel it first and hardest.
Yes at the top, shakier at the bottom. Organizations need people who can frame problems, design experiments, and judge whether the numbers mean anything more than ever — but the routine tier of querying, cleaning, and standard modeling is being automated by the field's own tools. Enter (or stay) with a plan to reach the judgment-heavy layer quickly, because the apprenticeship layer is thinning.
They replace tasks that used to fill most of the week: pipeline code, model search, tuning, and boilerplate analysis. What they don't replace is knowing which question to ask, detecting when data is lying to you, and convincing an organization to act on results. Teams are getting smaller per unit of output — that's the honest form the 'replacement' takes.
Causal inference and experiment design, since correlation-drawing is commoditized; deep fluency with LLM tooling, since directing AI is the new baseline; domain expertise in your industry; and communication strong enough to change decisions. Statistical judgment about what can go wrong — bias, leakage, confounding — is the durable core, because AutoML is confidently wrong in exactly those places.
They're contracting and changing shape. The classic first job — cleaning data and building basic models under supervision — is precisely what AI assistants now do, so employers hire fewer juniors and expect more from them. Breaking in increasingly requires demonstrated end-to-end work: a real problem framed, analyzed, and communicated, not a portfolio of tutorial models.