Manager, AI Solutions
United States · Remote · Permanent
Heads up: this posting is for future opportunities rather than one specific open role. If you apply, we'll add you to our candidate network and may reach out when relevant roles come up.
Axial Search is a specialist executive search firm built for one kind of hire: leaders who help organizations navigate AI transformation. Apply today to express your interest in roles like this one.
Visit our website to learn more about our process and explore free tools for your job search, including our live job market dashboard with salary, skills and hiring trend data from thousands of AI transformation roles.
What the market looks like
We've tracked 69 management-level MLOps postings across the US in the last six months, concentrated in California, Texas, and New York. Demand spans technology, financial services, healthcare, and manufacturing sectors. Compensation for this cohort typically ranges from $200K–$320K, reflecting the technical depth and operational ownership the role requires. The strongest candidates bring hands-on experience shipping and scaling ML infrastructure—they've built data pipelines, containerization strategies, and monitoring systems, and they can translate that execution experience into team leadership and cross-functional strategy.
Job responsibilities
Own the design, deployment, and lifecycle management of ML infrastructure and platform services that support model training, experimentation, and production inference at scale
Lead and mentor a team of MLOps engineers and platform engineers, setting technical direction, unblocking execution, and fostering a culture of operational excellence
Partner with data science, ML engineering, and software engineering teams to define SLAs, reliability standards, and tooling that accelerate model delivery and reduce time-to-production
Drive standardization around model versioning, experiment tracking, containerization, and CI/CD pipelines; build or integrate tools that reduce friction in the ML development workflow
Own observability, monitoring, and incident response for production ML systems; establish practices for detecting model drift, data quality issues, and performance degradation
Build and maintain cost optimization strategies around compute infrastructure, data storage, and cloud resources; track and report on platform utilization and efficiency metrics
Collaborate with security and governance teams to embed compliance, data privacy, and audit logging into platform design
Candidate requirements
6+ years of hands-on experience in MLOps, platform engineering, or ML infrastructure roles; at least 2–3 years managing or mentoring a technical team
Demonstrated proficiency building and scaling ML pipelines, containerization (Docker, Kubernetes), orchestration (Airflow, Kubeflow, or equivalent), and CI/CD systems for model deployment
Deep familiarity with cloud platforms (AWS, Azure, or GCP) and experience designing cost-efficient, reliable ML infrastructure on cloud services
Strong understanding of software engineering practices and ability to write production-quality code; comfort working in Python, Go, or similar languages
Experience defining and monitoring SLAs for production systems; background with observability tools, logging, and incident response in ML or data-intensive environments
Proven ability to communicate technical trade-offs to non-technical stakeholders; track record of aligning infrastructure decisions with business outcomes