ML engineering & MLOps

Turning notebooks into training pipelines, model registries and deployments that survive a second engineer.

  • Python
  • PyTorch
  • MLflow
  • Airflow
  • Dagster
  • Ray
  • Docker
  • Kubernetes

Research code and production code answer different questions. We take the model that already works on your researcher’s laptop and give it reproducible training runs, versioned data and artefacts, a registry, and a deployment path that does not depend on anyone remembering the incantation.

The first question we ask is whether you can retrain last month’s model today — exactly, not roughly. If the answer is no, everything downstream is guesswork: you cannot attribute a regression, roll back with confidence, or reproduce a production incident. Fixing that is usually cheap and almost always first.

Then we write down how it works, because the point is that your team owns it afterwards. A pipeline only one consultant understands is a worse outcome than the notebook it replaced.

What you get

  • Reproducible training pipelines with tracked runs and pinned data versions
  • Model registry, promotion rules and rollback
  • Batch and real-time inference services with autoscaling
  • Drift, latency and cost monitoring with alerts that fire before customers notice

Common questions

Do you build models, or only the engineering around them?
Mostly the engineering around them, and we say so deliberately. If you have a data scientist whose model works, we are the fastest route to it running in production. If you need novel modelling research, we are not the right team and will tell you on the first call.
We are on a single big training script. Where do you start?
Reproducibility, then decomposition. Pin the dependencies and data, get the same run producing the same weights twice, and only then split it into steps. Reordering that sequence is how teams end up with a pretty pipeline that produces a different model every Tuesday.
Can you cut our GPU bill?
Often, yes — through batching, right-sizing, spot capacity with checkpointing, and killing training runs nobody reads the output of. We quote that work against measured spend, so you can see whether it pays for itself before committing.

Next step

Tell us what has to ship, and by when.

One call, no deck. If we are not the right team for it we will say so, and usually point you at who is.

Start a conversation branislav@thrivee.io

Typical reply within one business day · CET / CEST (UTC+1 / UTC+2)