Principal Product Manager, AI Infrastructure, Orchestration

Posted yesterday

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Principal Product Manager owning DataRobot’s Kubernetes orchestration platform for AI agents and models. Defining placement, scaling, lifecycle, isolation, APIs, governance, and reliability across enterprise deployments.

Responsibilities

  • Own the deployment and workload API, including the resource model, lifecycle semantics, versioning, backward compatibility, and customer-facing error behavior
  • Define placement and capacity across nodes and accelerators, including quota, priority, and contention handling across tenants
  • Own autoscaling signals, cold-start and scale-to-zero economics, headroom policy, and cost-versus-latency trade-offs
  • Own agent runtime decisions involving execution location, duration, isolation, tool calls, and state persistence through restarts or evictions
  • Define ingress, routing, request-aware load balancing, tenancy boundaries, and private connectivity for model and agent endpoints
  • Own governance and audit capabilities covering deployments, invocations, policies, access control, and audit durability
  • Define metering, quota, tenant attribution, inference packaging, and pricing
  • Own reliability SLOs, error budgets, and operational diagnostics for degraded deployments
  • Hold the roadmap for the workload and agent runtime layer across two engineering pods
  • Align with product teams building on top of the platform
  • Operate through design reviews, API contracts, and production data

Requirements

  • 6+ years in product management for infrastructure, developer platforms, or cloud services
  • At least 3 years working on Kubernetes-based or distributed systems products
  • Principal candidates bring 9+ years of experience and have owned a platform layer used by other product teams
  • Deep technical understanding of GPU and accelerator behavior, including topology-aware placement, fractional and time-sliced sharing, MIG, device plugins, driver and container-runtime plumbing, memory constraints, and utilization economics
  • Deep technical understanding of Kubernetes, including the API server and scheduler, controllers and CRDs, operators, admission and RBAC, device plugins, resource requests and limits, node pools, and pod scheduling failures
  • Multi-tenancy experience, including isolation models, noisy neighbors, quota and fairness, and security-review-ready tenancy designs
  • Experience owning a public or platform API
  • Technical writing and prototyping ability, including documentation, deep dives, public posts, API references, or working prototypes
  • Comfort operating with matrixed engineering teams and no direct reports
  • BS or MS in Computer Science or a closely related technical field, or equivalent hands-on experience as a software, platform, or infrastructure engineer
  • Nice to have: service networking depth, modern serving stacks such as vLLM, long-running and agentic workload patterns, customer-managed/air-gapped/sovereign deployments, regulated-industry experience, and CNCF or open-source contribution
  • Ability to build prototypes, evaluation harnesses, agents connected to real services, or other working tools independently

Benefits

  • Medical, Dental & Vision Insurance
  • Flexible Time Off Program
  • Paid Holidays
  • Paid Parental Leave
  • Global Employee Assistance Program (EAP)
  • Competitive pay
  • Reasonable accommodations for applicants with physical and mental disabilities

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

CloudDistributed SystemsKubernetesNode.js

Location requirements

RemoteUnited States

Report this job

Found something wrong with the page? Please let us know by submitting a report below.