DevOps and Machine Learning Operations Engineer
Location
Manchester, UK
Contract Type
Contract position
About the Project
Join a specialist software development team delivering a new, business-critical technology platform for an established international organisation. This greenfield development project involves modern cloud architecture, data-intensive applications, and AI-enabled capabilities. You will work as part of a multidisciplinary project team alongside experienced software, AI, infrastructure, and design professionals, with direct involvement in taking the platform from development through to production. Security, scalability, maintainability, data protection, and production readiness are key priorities throughout the project.
Role Purpose
The DevOps and MLOps Engineer will build and operate the delivery pipelines, environments, and runtime platform the project depends on, together with the model-serving path. The role covers infrastructure as code, continuous integration and deployment, observability, cost control, and production support. Applicants must have experience running systems in production and being accountable for their reliability.
Main Responsibilities
- Infrastructure and Environments:
- Define and maintain cloud infrastructure as code, ensuring reviewable changes and reproducible environments for development, staging, and production.
- Manage secrets, certificates, network boundaries, and access with least privilege as the default and auditable grants.
- Maintain consistency across environments and make explicit any necessary differences.
- Continuous Integration and Deployment:
- Build pipelines that test, scan, build, and deploy, with quality gates that fail closed.
- Automate database migration and rollback for reversible releases.
- Support progressive delivery, including staged rollout and fast rollback, with deployment records.
- Observability and Operations:
- Instrument services with structured logging, metrics, and tracing; define alerts based on user-impacting symptoms.
- Establish service level objectives and report against them.
- Run incident response: triage, mitigate, restore, and document post-incident reviews with tracked corrective actions.
- Machine Learning Operations:
- Package, version, and deploy models and dependencies for traceable predictions.
- Automate evaluation before promotion and monitor drift, latency, and cost in production.
- Enable safe rollback of models independently of application releases.
- Manage inference cost and capacity, including batching, caching, and hardware selection.
- Security and Compliance:
- Apply dependency and container scanning, patching, and image provenance as part of the pipeline.
- Support data protection obligations, including retention, deletion, and access logging.
- Document and rehearse recovery objectives.
Essential Technical Experience
- Substantial professional experience in DevOps, platform, or site reliability engineering on production systems with accountability for reliability.
- Strong infrastructure as code practice (e.g., Terraform) and experience with a container orchestration platform such as Kubernetes.
- Demonstrated CI/CD ownership, including build, test, security gates, deployment, and rollback.
- Practical observability experience: structured logging, metrics, tracing, alert design, and on-call.
- Experience deploying and operating machine learning models in production, including versioning, evaluation, and monitoring.
- Competence in at least one of Python, Go, or TypeScript for automation and tooling.
- Sound understanding of cloud networking, identity and access management, and secret handling.
- Ability to reason about cost and explain architectural trade-offs to both engineers and business stakeholders.
Desirable Experience
- Experience with GPU scheduling and inference optimisation.
- Experience with feature stores, vector databases, or retrieval pipelines.
- Exposure to regulated environments and formal audit.