Senior Support Engineer

Posted 6 hours ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Senior Support Engineer troubleshooting OpenAI API infrastructure for strategic enterprise customers. Designing observability, incident response, and automation processes for reliable AI deployments.

Responsibilities

  • Serve as a foremost technical and troubleshooting expert for OpenAI’s API platform and act as the last line of defense before core Engineering
  • Identify and implement scalable support operations using automation and AI technologies
  • Configure and use monitoring and alerting workflows to detect customer-impacting issues in real time
  • Partner with Engineering on reliability reviews and operational readiness for features, launches, and strategic customer requirements
  • Design and refine incident response processes and documentation
  • Analyze operational metrics and incident root-cause analyses to improve monitoring dashboards, alert configurations, and support workflows
  • Provide support coverage during holidays and weekends based on business needs
  • Collaborate directly with strategic enterprise customers, product teams, Infrastructure, Engineering, Technical Success, and Support teams

Requirements

  • Bachelor’s degree in Computer Science or a related field
  • 8+ years of experience in technical operations roles such as SRE/NOC
  • Experience designing monitoring systems and resolving production issues in fast-paced, mission-critical environments
  • Strong track record troubleshooting complex technical problems at the systems level
  • Deep familiarity with monitoring, alerting, and observability practices
  • Hands-on experience with metrics, logging, and tracing for distributed systems, including SLIs/SLOs, alert tuning, and dashboard creation
  • Experience leading incident response for high-severity outages or service disruptions
  • Ability to coordinate incidents in real time, perform root cause analysis, and drive post-mortems and action items
  • Knowledge of industry best practices for incident management and fault diagnosis
  • Strong scripting or software engineering skills, such as Python or similar
  • Understanding of cloud infrastructure and distributed systems fundamentals
  • Familiarity with cloud services, load balancers, databases, and containerized applications
  • Ability to work cross-functionally and explain technical issues to engineering and non-technical stakeholders
  • Ability to provide support coverage during holidays and weekends based on business needs

Benefits

  • Remote work arrangement
  • Reasonable accommodations for applicants with disabilities
  • Immigration and sponsorship support may be provided based on individual circumstances

Job type

Full Time

Experience level

Senior

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

CloudDistributed SystemsPython

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.