Staff AI Platform Engineer – Inference, Agentic Systems

Posted 3 weeks ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff AI Platform Engineer building Paytm Labs' inference and agentic AI platforms. Operating models and autonomous workflows across payments, risk, fraud, and enterprise fintech systems.

Responsibilities

  • Build and operate multi-model serving across text, voice, code, and vision modalities
  • Own the model lifecycle: download, deploy, serve, monitor, update, and swap models
  • Optimize inference latency, throughput, and cost through quantization, batching, caching, and routing strategies
  • Ensure inference reliability for agents and dependent systems
  • Architect and build the Agentic AI Platform, including runtime infrastructure, orchestration systems, and developer tooling
  • Design multi-agent coordination systems for complex workflows
  • Build secure tool-use infrastructure for APIs, databases, and services
  • Implement guarded workflow automation for multi-step business and engineering tasks
  • Build safety and guardrail systems, including permissioning, sandboxing, and human-in-the-loop workflows
  • Develop evaluation and observability frameworks for agent behavior, regression detection, and failure debugging
  • Develop SDKs and APIs for internal teams to build and deploy agents
  • Define technical direction and architecture for agentic systems across the organization
  • Establish patterns and standards for agent design, tool calling, and evaluation
  • Partner with ML, product, and security teams
  • Mentor engineers and contribute to agent system design best practices

Requirements

  • 8+ years of software engineering experience, with 3+ years in AI systems or LLM applications
  • Strong understanding of LLM-based agent architectures: tool use, multi-step workflows, multi-agent coordination, and failure modes
  • Experience building highly reliable distributed systems
  • Experience evaluating LLM systems in production, including building evals, detecting regressions, and debugging non-deterministic failures
  • Proficiency in TypeScript or Python, and willingness to work in both
  • Experience with modern LLM APIs or open-source models
  • Experience with or strong interest in model serving, including vLLM, TensorRT-LLM, or Triton
  • Understanding of distributed systems, including task queues, event-driven architectures, state management, and durable long-running workflows
  • Experience with AWS or GCP and containerized deployments
  • Strong understanding of security risks in agentic systems, including prompt injection, privilege escalation, and data leakage
  • Demonstrated experience leading complex technical initiatives
  • Strong written and verbal communication skills

Benefits

  • Full-time employment
  • Diversity and equal opportunity commitment
  • Accessibility accommodations during recruitment and selection
  • Inclusive, discrimination-free workplace commitment

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

AWSDistributed SystemsGoogle Cloud PlatformPythonTypeScript

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.