■ HIGH RISK ■ Technology
Roughly half of it, on current trajectory. Pipeline authoring is being automated aggressively, but data engineering has always been less about writing transformations than about the fact that upstream systems lie.
“Pipelines and ETL are getting automated fast. But data is still messy, and someone has to clean up the mess.”
Our AI replacement risk score — how we score jobs
The job is moving and reshaping data reliably: ingesting from operational databases, SaaS APIs and event streams, landing it in a warehouse or lakehouse, modelling it into tables analysts can trust, orchestrating the whole thing on a schedule, and monitoring for the morning when a dashboard reports zero revenue. Add governance work such as access controls, personally identifiable data handling, lineage documentation and retention policy, plus cost management, since a badly written query on a cloud warehouse can cost more than the engineer who wrote it.
Generation and managed tooling are absorbing large parts of this. Connectors that once took a week are now off-the-shelf, models write SQL transformations and orchestration configuration competently, schema mapping and documentation generate well, and modern platforms handle scheduling, retries and scaling automatically. Data quality tooling detects anomalies without bespoke code, and natural-language querying is chipping away at the demand for hand-built reporting tables. Our score of 55 reflects a discipline whose most visible artefacts, pipelines and SQL, are highly generatable.
What resists is everything caused by organisations rather than by code. Source systems change without warning, semantics drift, two departments define revenue differently, and a field that used to mean one thing quietly starts meaning another after a product launch. Resolving that requires knowing who to ask and having the standing to make a definition stick. Governance carries legal weight, particularly around personal data, and someone must be accountable for who can see what. Cost architecture, latency trade-offs between batch and streaming, and the design of a model that survives reorganisation are judgement calls. Data engineering is also currently absorbing new work from AI systems, which need retrieval corpora, embeddings pipelines and evaluation data managed with the same rigour, and that is buoying demand while the routine pipeline work erodes.
Automatability: our editorial assessment of current and near-term AI capability
Real pressure by around 2030. Managed connectors and generated transformations have already cut the hours needed for a standard warehouse build substantially, and natural-language analytics is reducing demand for bespoke reporting models. Expect teams to shrink on pipeline construction while shifting toward platform ownership, governance and AI data infrastructure. The engineers who feel it first are those whose job is writing connectors and dbt models against well-behaved sources.
The construction layer is. Connectors, transformations, orchestration config and documentation all generate well, and managed platforms handle infrastructure that used to require a team. What remains stubbornly manual is dealing with sources that change without notice, conflicting business definitions, governance obligations, and the debugging of failures that produce plausible but wrong numbers.
It reduces demand for hand-built reporting tables and simple analyst requests, which is a genuine shift. It also increases the importance of a trustworthy semantic layer, because a system answering questions in plain English against badly modelled data produces confident nonsense at scale. Somebody has to define the metrics it answers with, and that is data engineering work.
Data modelling that survives organisational change, semantic layer ownership, governance and privacy engineering, cost architecture on cloud warehouses, and reliability practice for pipelines that fail silently. Increasingly also the infrastructure behind AI features: document processing, embeddings pipelines, retrieval quality and the datasets used for evaluation.
Yes, with a clear eye about what the role is becoming. Entry-level work that consisted of building connectors and writing transformations is shrinking. The growth is in platform, governance and AI data infrastructure, which expect more engineering maturity. Coming in through software engineering or analytics and moving toward data platform ownership is a more reliable route than pipeline authoring alone.