MLOps & Orchestration

Cluster Operations Engineer

Boston Metro, MA · Hybrid or Remote · Full-time

Apply now→

A cluster that passes acceptance is the starting line, not the finish line.

Our customers get dedicated NVIDIA clusters and a named engineer who knows their workload, not a ticket queue. This is that engineer, for everything above the metal: Kubernetes and Slurm orchestration, the NVIDIA software stack, control plane health, and the customer relationship that turns "here's what we're training" into a cluster configured to do it well.

This role is for you if
  • →You've operated Kubernetes or Slurm in production and have opinions about scheduler policy.
  • →You want to run the NVIDIA toolchain (DCGM, Base Command Manager, NCCL) on the newest hardware NVIDIA makes.
  • →You like the customer side: onboarding an ML team, learning what they're building, and shaping the platform around it.
  • →You believe a control plane should be boring, monitored, and recoverable.
What we promise you
  • →Direct ownership, with no layer between you and the clusters or the customers on them.
  • →A fleet spanning GB300 NVL72 to H100, on Quantum-3 XDR InfiniBand, scaling toward 9,552 GPUs.
  • →Customers whose training runs you'll know by name.

At a glance

Team
MLOps & Orchestration
Location
Boston Metro, MA
Arrangement
Hybrid or Remote
Apply now→

Or write to careers@cometcloud.com

Other open roles