Senior Platform Engineer developing infrastructure for SaaS product at Orchestry. Responsible for building production-grade, containerized systems and ensuring scalability and reliability in the platform.
Responsibilities
Convert existing development-only Dockerfiles into production-ready, multi-stage container builds with health checks and graceful shutdown handling (including draining in-flight background jobs before termination)
Stand up and administer a container registry, including access policies, image tagging conventions, and retention/cleanup automation
Extend CI/CD pipelines to add container build-and-push steps, and migrate deployment targets from traditional package-based deployment to container-based App Service deployments
Author Infrastructure-as-Code modules from scratch covering compute, database, storage, secrets, and monitoring resources
Build provisioning automation for identity/app-registration setup, database initialization, and environment bootstrapping
Implement liveness and readiness health endpoints that verify downstream dependency connectivity (database, queue, background job processing)
Build a telemetry pipeline that exports sanitized, PII-free operational metrics to a centralized monitoring system, with configurable scope and destination
Design and maintain a tracking/registry system for infrastructure deployments, capturing version, status, and configuration state across many environments
Extend deployment pipelines to support automated, tag-based promotion and staged/canary rollout strategies, with per-environment failure isolation and rollback
Build a controlled mechanism for pushing critical patches outside the normal release cadence, with audit logging, notification workflows, and approval gates
Partner with engineering to audit core subsystems (background job processing, database partitioning/sharding, secrets access patterns, external integrations, feature-flag and analytics tooling, licensing/validation logic, notifications, scheduled tasks) for portability across deployment environments
Optimize databases (Azure SQL, Cosmos DB) for performance and scalability, including sharding and partitioning strategies
Ensure horizontal scaling strategies to handle SaaS growth efficiently
Implement caching (Redis, CDN) and performance tuning techniques
Lead incident response and root cause analysis (RCA) efforts, reducing mean time to recovery (MTTR)
Participate in the on-call rotation as a first responder to production incidents and emergency scenarios, providing timely triage and resolution outside of standard business hours
Work closely with Engineering, Security, and Product teams to align platform goals with business objectives
Act as a technical mentor for junior and mid-level engineers, fostering best practices in DevOps, cloud, and automation
Senior Platform Engineer building scalable backend and cloud infrastructure for ExaCare’s AI - powered post - acute care platform. Improving reliability, developer velocity, and healthcare admission workflows.
Principal Platform Engineer leading SkyWatch’s satellite - data platform and AI agent infrastructure. Owning architecture, customer - driven roadmap delivery, and platform engineering leadership.
Platform Engineer securing Just Eat Takeaway.com’s global food - delivery edge infrastructure. Building gateways, automation, and resilient traffic routing across production environments.
Director leading Blackpoint Cyber’s cloud - based Unified Security Posture data platform for cybersecurity solutions. Driving platform roadmap, reliability, APIs, data engineering, and team growth.
Ingénieur logiciel principal intégrant des plateformes, API et solutions IA chez EDC. Gouvernance technique, sécurité, résilience et mentorat dans une société canadienne de financement du commerce.
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.