Senior DevOps/MLOps Engineer

Remote Senior MLOps Engineer

Interested in this MLOps Engineer role at Leverege?

Apply Now →

Skills & Technologies

AwsAzureGcpKubernetesPython

About This Role

AI job market dashboard showing open roles by category

#### Engineering

Senior DevOps/MLOps Engineer

================================

Remote

Full\-Time

Job Description

-------------------

Elevator Pitch

------------------

Leverege is hiring a Senior DevOps/MLOps Engineer. We build AI\-native software that turns cameras into real\-time visibility into physical operations, and our computer vision runs everywhere from managed Kubernetes in the cloud to a growing fleet of on\-site edge servers in auto service centers, factories, and stores. This role owns the infrastructure that keeps all of it running and builds the systems that let us deploy and manage that edge fleet at scale.

This is a rare chance to join a high\-trust, fully remote company and own platform and MLOps work that directly moves the business. If you love Kubernetes, think in terms of fleets rather than single servers, and want to own the GPU and edge infrastructure that gets computer vision to customers reliably, we would love to hear from you!

The Opportunity

-------------------

Leverege builds VisionAI software that helps businesses see what is happening across their physical operations. We connect to cameras, run proprietary computer vision models (often on a small edge appliance on\-site), and deliver real\-time operational insights to Fortune 500 customers across automotive service, manufacturing, and retail. Our products run on the Leverege Stack across per\-customer environments in Google Cloud, with GPU and inference infrastructure supporting our machine learning team and a fleet of edge servers deployed at customer locations.

That footprint is scaling fast, from dozens of deployments toward hundreds, and the infrastructure needs to scale with it. Today, a lot of edge server provisioning and management is manual. We need someone to own the platform, harden it, and build the automation that makes deploying and updating a large edge fleet routine and safe.

You will be a senior individual contributor on the DevOps team. Success means our products ship reliably, our GPU and inference infrastructure keeps up with the ML team, incidents are rare and quickly resolved, and managing hundreds of edge servers feels as controlled as managing one.

What You’ll Do

------------------

You will own the infrastructure that runs Leverege’s VisionAI in production, in the cloud and at the edge, and build the systems that let us operate it at scale.

  • Own and operate the GKE platform across our per\-customer Google Cloud projects, including provisioning new customer clusters end to end with Terraform, Helm, and GitOps.
  • Build the edge fleet management systems that let us deploy, monitor, update, and roll back software across a growing fleet of on\-site edge servers, replacing manual per\-server work with automated, auditable processes.
  • Run the MLOps path for computer vision by building repeatable pipelines that get GPU workloads and CV models from the ML team onto inference nodes and the edge fleet reliably.
  • Keep production healthy and observable with strong instrumentation and alerting (Prometheus, Grafana, Sentry), and serve as a primary responder on the on\-call rotation.
  • Optimize cost and capacity across GKE and GPU node pools, balancing spend against reliability as the deployment count grows.
  • Harden security and support compliance through least\-privilege access, proper secrets management, and SOC 2 evidence for the infrastructure you own.

Who You Are

---------------

  • A driver, not a passenger. You take ownership of production and edge systems end to end and act before being asked. No one will look over your shoulder, and you prefer it that way.
  • A fleet thinker. You instinctively design for many machines across many environments, not one server at a time, and you plan for intermittent connectivity, remote updates, and things going wrong far from your keyboard.
  • Calm and methodical under pressure. When production breaks or an edge site goes dark, you debug systematically and communicate clearly rather than thrashing.
  • Direct and collaborative. You push back on risky changes, escalate straight to the right person regardless of rank, and tell product and ML teams honestly when something is not safe to ship, while staying kind and easy to work with async.
  • An automator by instinct. You would rather build the tool once than do the toil forever, and you improve shared tooling so the whole team moves faster.
  • Comfortable with autonomy and ambiguity. This is a fast\-scaling environment where not everything is documented yet. If you need a lot of structure and hand\-holding, this is not the right fit.

Qualifications (Required)

-----------------------------

  • 8\+ years in DevOps, SRE, platform, or infrastructure engineering, with real production ownership.
  • Expert\-level Kubernetes in production (managed Kubernetes such as GKE strongly preferred).
  • Strong Infrastructure\-as\-Code experience with Terraform and Helm.
  • Deep experience on a major cloud provider (GCP preferred; strong AWS or Azure background with willingness to work primarily in GCP is fine).
  • Experience operating a distributed fleet of remote or edge servers, or comparable experience managing infrastructure across many isolated environments.
  • Hands\-on experience running GPU workloads and/or deploying ML models to production (inference serving, model rollout).
  • Solid CI/CD, Linux, networking, and scripting (Python, Node, Go, or Bash) fundamentals.
  • Experience as a primary on\-call responder for production systems.

Qualifications (Preferred)

------------------------------

  • Experience with model\-serving stacks (Triton) and computer vision or ML data pipelines.
  • GitOps with ArgoCD, and observability with the Prometheus operator and Grafana.
  • Operating stateful services in Kubernetes (CloudNativePG/PostgreSQL, Elasticsearch, Redis) and event streaming (Pub/Sub).
  • Networking and connectivity for distributed fleets (Tailscale, Cloudflare) and camera or video streaming such as RTSP.
  • SOC 2 or similar compliance experience, and familiarity with GCP access controls (IAM, PAM, Secret Manager, External Secrets).
  • Exposure to edge hardware and on\-site compute in real deployments.

Why Leverege

----------------

  • Computer vision applied to the physical world. We build proprietary vision models that solve operational problems for businesses in automotive, manufacturing, and retail, running on infrastructure you will own.
  • Infrastructure that is core to the business. The edge fleet and platform are how the product reaches customers, so your work is visible and high\-impact.
  • Clear product\-market fit with room to run. Fortune 500 customers, a growing portfolio of products, and large markets with little direct competition.
  • Fully remote, permanently. Work from anywhere in the US. We have committed to remote for good and have figured out how to make it work.
  • High\-trust culture without politics. Smart, kind, driven people. High expectations, but the hard part is the problems, not internal friction.
  • A company scaling from startup to mid\-size. Real opportunity to shape how our infrastructure and MLOps practices grow as we scale toward hundreds of deployments.

Role Details

Company Leverege
Title Senior DevOps/MLOps Engineer
Location Remote, US
Category MLOps Engineer
Experience Senior
Salary Not disclosed
Remote Yes

About This Role

MLOps Engineers build the infrastructure that keeps ML models running in production. They own CI/CD pipelines for model deployment, monitoring for data drift and model degradation, and the tooling that lets data scientists ship faster. If ML Engineers build the models, MLOps Engineers build the roads those models travel on.

The job is fundamentally about reliability and velocity. Data scientists want to iterate fast. Product teams want stable predictions. Your job is to make both happen simultaneously. That means building deployment pipelines that catch regressions before they hit production, monitoring systems that alert on data drift before it degrades model performance, and self-service tooling that lets data scientists deploy without filing a ticket.

Across the 3,708 AI roles we're tracking, MLOps Engineer positions make up 1% of the market. At Leverege, this role fits into their broader AI and engineering organization.

MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.

What the Work Looks Like

A typical week involves: debugging a model deployment that's serving stale predictions, building a new monitoring dashboard for a feature team, writing Terraform for GPU-enabled inference clusters, reviewing pull requests for the ML platform's CI/CD pipeline, and meeting with data scientists to understand their pain points. You're the bridge between ML and infrastructure.

MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.

Skills Required

Aws (30% of roles) Azure (24% of roles) Gcp (17% of roles) Kubernetes (12% of roles) Python (51% of roles)

Kubernetes, Docker, and cloud infrastructure are baseline. Most roles want experience with ML-specific tooling: MLflow, Kubeflow, Weights & Biases, or similar. Strong DevOps fundamentals matter more than ML theory. You need to understand model serving (TorchServe, Triton, vLLM), monitoring (Prometheus, Grafana), and infrastructure-as-code (Terraform, Pulumi).

GPU infrastructure knowledge is increasingly valuable as LLM inference becomes a major cost center. Understanding GPU scheduling, multi-node training setups, and inference optimization (quantization, batching, caching) puts you in the top tier. Experience with model registries and feature stores rounds out the profile.

Good MLOps postings specify their ML stack, infrastructure scale, and the problems they're solving (deployment velocity, cost optimization, monitoring gaps). Red flag: companies that want MLOps but don't have any models in production yet. You'll end up doing general DevOps instead.

Compensation Benchmarks

MLOps Engineer roles pay a median of $220,000 based on 47 positions with disclosed compensation. Senior-level AI roles across all categories have a median of $230,000.

Across all AI roles, the market median is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. For comparison, the highest-paying categories include AI Safety ($300,000) and Research Engineer ($280,000). By seniority level: Entry: $120,000; Mid: $200,000; Senior: $230,000; Director: $272,150; VP: $250,000.

Leverege AI Hiring

Leverege has 1 open AI role right now. They're hiring across MLOps Engineer. Based in Remote, US.

Remote Work Context

Remote AI roles pay a median of $185,334 across 717 positions. About 14% of all AI roles offer remote work.

Career Path

Common paths into MLOps Engineer roles include DevOps Engineer, Platform Engineer, Data Engineer.

From here, career progression typically leads toward ML Platform Lead, Infrastructure Architect, Engineering Manager.

DevOps engineers with ML curiosity have the shortest path. You already understand deployment, monitoring, and infrastructure. Add ML-specific knowledge (model serving, data pipelines, experiment tracking) and you're competitive. The career ceiling is high: ML Platform Lead roles at top companies pay well because the infrastructure complexity is enormous.

What to Expect in Interviews

Interviews emphasize infrastructure and reliability. Expect questions about CI/CD for ML models, monitoring for data drift, and how you'd design a model serving platform that handles 10K requests per second. Coding rounds focus on Python and infrastructure-as-code (Terraform, Helm). Be ready to discuss tradeoffs between different model serving frameworks and how you'd handle rollback when a new model degrades performance.

When evaluating opportunities: Good MLOps postings specify their ML stack, infrastructure scale, and the problems they're solving (deployment velocity, cost optimization, monitoring gaps). Red flag: companies that want MLOps but don't have any models in production yet. You'll end up doing general DevOps instead.

AI Hiring Overview

The AI job market has 3,708 open positions tracked in our dataset. By seniority: 102 entry-level, 1,705 mid-level, 1,469 senior, and 432 leadership roles (Director, VP, C-Level). Remote roles make up 14% of the market (508 positions). The remaining 3,180 roles require on-site or hybrid attendance.

The market median for AI roles is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. Highest-paying categories: AI Safety ($300,000 median, 21 roles); Research Engineer ($280,000 median, 147 roles); AI Architect ($254,798 median, 67 roles).

MLOps demand tracks closely with production ML adoption. As more companies move models from notebooks to production, the need for MLOps grows. The role is well-established at large tech companies and growing fast at mid-stage startups that are hitting the 'our models work in notebooks but break in production' phase.

The AI Job Market Today

The AI job market spans 3,708 open positions across 16 role categories. The largest categories by volume: AI/ML Engineer (2,605), Data Scientist (310), AI Software Engineer (259). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.

The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (102) are outnumbered by mid-level (1,705) and senior (1,469) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 432 positions, representing the bottleneck between technical execution and organizational strategy.

Remote work availability sits at 14% of all AI roles (508 positions), with 3,180 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.

AI compensation is structured in clear tiers. The market median sits at $217,500. Top-quartile roles start at $272,100, and the 90th percentile reaches $325,000. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.

Category matters for compensation. AI Safety roles lead at $300,000 median, while Prompt Engineer roles sit at $140,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.

The most in-demand skills across all AI postings: Python (1,890 postings), Aws (1,103 postings), Azure (877 postings), Rag (855 postings), Gcp (631 postings), Prompt Engineering (560 postings), Pytorch (545 postings), Claude (498 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.

Frequently Asked Questions

Based on 47 roles with disclosed compensation, the median salary for MLOps Engineer positions is $220,000. Actual compensation varies by seniority, location, and company stage.
Kubernetes, Docker, and cloud infrastructure are baseline. Most roles want experience with ML-specific tooling: MLflow, Kubeflow, Weights & Biases, or similar. Strong DevOps fundamentals matter more than ML theory. You need to understand model serving (TorchServe, Triton, vLLM), monitoring (Prometheus, Grafana), and infrastructure-as-code (Terraform, Pulumi).
About 14% of the 3,708 AI roles we track offer remote work. Remote availability varies by company and seniority level, with senior and leadership roles more likely to offer location flexibility.
Leverege is among the companies actively hiring for AI and ML talent. Check our company profiles for detailed breakdowns of open roles, salary ranges, and hiring trends.
Common next steps from MLOps Engineer positions include ML Platform Lead, Infrastructure Architect, Engineering Manager. Progression depends on whether you lean toward technical depth, people management, or product strategy.

Get Weekly AI Career Intelligence

Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.