Most coverage of AI focuses on the tools people use, such as chatbots, copilots and image generators. Far less attention goes to the engineers who build the systems underneath them. AI and machine learning engineers are the people who take a model that works in a notebook and turn it into something a business can depend on every day.
Calling them architects is not flattery. The hard part of production AI is not training a model once; it is designing the surrounding system so the model keeps getting the right data, keeps performing as the world changes, and fails safely when it does not. This article looks at that architecture layer by layer, and at what it means for organizations that want AI to be more than a demo.
The model is the smallest part of the system
In a widely cited 2015 paper, “Hidden Technical Debt in Machine Learning Systems,” a team of Google engineers observed that in real-world ML systems only a small fraction of the code is the model itself. The rest is data collection, data verification, feature extraction, configuration, serving infrastructure, monitoring and process management.
That observation has aged well. A fraud model, a demand forecast or a document classifier depends on a chain of systems, and each link can break quietly. An upstream team renames a field. A vendor changes a file format. Customer behavior shifts after a price change. None of these throw an error message; the model just starts giving worse answers.
An ML engineer’s job is to design that chain so problems are visible and recoverable. That is the difference between a data scientist’s prototype and an engineered system.
The five layers an ML engineer designs
It helps to think of a production ML system as five layers, each with its own design decisions.
Data. Where training and live data come from, how they are validated, and how the features the model sees in production are kept identical to the ones it was trained on. Mismatches between training data and live data, often called training-serving skew, are one of the most common causes of models that test well and perform badly.
Training. Reproducible pipelines that can retrain a model on demand, with versioned data, code and parameters, so any model in production can be traced back to exactly how it was built. Tools such as MLflow, Kubeflow and TFX exist largely to make this repeatable.
Serving. How predictions are delivered: in real time behind an API, in nightly batches, or on a device at the edge. Each option has different costs, latency and failure modes. A recommendation that arrives two seconds late may be worthless; a nightly risk score may not need to be instant at all.
Monitoring. Tracking not just whether the service is up, but whether its inputs and outputs still look like they did when the model was validated. This is where drift is caught.
Feedback. Capturing what actually happened, such as whether the flagged transaction really was fraud, and feeding it back into evaluation and retraining. Without a feedback loop, a model cannot improve and nobody can prove it is still accurate.
Why models decay, and how engineers plan for it
Traditional software does the same thing tomorrow that it did today. Machine learning models do not, because the world they learned from keeps moving. Economists call a similar effect regime change; ML engineers call it drift.
There are two broad kinds. Data drift is when the inputs change: a new product line, a new customer segment, a different mix of devices. Concept drift is when the relationship between inputs and outcomes changes: fraudsters adopt new tactics, or a policy change alters what “normal” looks like. Both degrade accuracy without any code changing.
Engineers plan for this with MLOps practices borrowed from DevOps: automated retraining on a schedule or trigger, champion-challenger testing where a new model runs alongside the current one before replacing it, staged rollouts, and a fast path to roll back. The goal is to make model updates as routine and low-risk as software releases.
The skill stack behind the architecture
Because the job spans all five layers, the skill stack is broad. The core tools map neatly onto the architecture:
- Modeling: Python with PyTorch, TensorFlow and scikit-learn, plus working depth in at least one area such as deep learning, natural language processing or computer vision.
- Data pipelines: orchestration and streaming tools such as Apache Airflow and Apache Kafka, and solid SQL.
- Infrastructure: cloud platforms (AWS, Google Cloud or Azure), containers and infrastructure as code.
- MLOps: experiment tracking, model registries and pipelines with tools such as MLflow, Kubeflow or TFX.
- Judgment: choosing a simple, explainable model when it is good enough, and knowing when a decision tree beats a neural network.
The strongest engineers also think about security and responsibility: who can access training data, whether a model could leak sensitive information, and how its decisions can be explained to a regulator or customer.
What this means for businesses building AI
If you are commissioning or building AI, the architecture view changes what you should ask for. A few practical steps:
- Budget for the system, not the model. Plan time and money for data pipelines, monitoring and retraining from the start, not as a phase two that never arrives.
- Define how you will know it is working. Agree on business and model metrics, and on who reviews them and how often.
- Insist on reproducibility. Any model in production should be traceable to its data, code and configuration.
- Plan the feedback loop. Decide how real outcomes will be captured and fed back before launch.
- Decide who owns it after launch. A model without an owner will drift until someone notices a problem the hard way.
For many organizations, especially those using large language models from a vendor rather than training their own, the same principles apply to prompts, retrieval pipelines and evaluation sets. Our article on retrieval-augmented generation covers that pattern, and data flow development goes deeper on the pipeline skills involved.
Frequently asked questions
How is an ML engineer different from a data scientist?
A data scientist typically explores data, tests hypotheses and builds candidate models. An ML engineer focuses on turning those models into reliable production systems: pipelines, serving, monitoring and retraining. In small teams one person often does both, but the mindsets differ.
Do we need ML engineers if we only use vendor AI models?
You may not need people who train models from scratch, but you still need the engineering discipline: integrating the model with your data, evaluating output quality, monitoring cost and behavior, and managing updates when the vendor changes the model.
How often should a production model be retrained?
There is no universal schedule. It depends on how quickly your data changes. The better approach is to monitor for drift and performance decline and retrain when those signals cross an agreed threshold, with a scheduled retrain as a backstop.
Build AI that holds up in production
Delana Technologies helps businesses design the data, integration and monitoring architecture that turns models into dependable systems, whether you are training your own or building on vendor AI. Explore our AI consulting and agentic AI solutions, then call 239.414.5126 or contact us to talk through your project.
Sources: D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” NeurIPS 2015; documentation for MLflow, Kubeflow and TensorFlow Extended (TFX).
