Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Responsibilities
Provide production support and maintenance for enterprise applications hosted on Google App Engine
Triage, diagnose, restore service, and drive user-impacting incidents to closure within agreed SLAs/SLOs
Troubleshoot application errors, latency, performance degradation, service failures, configuration issues, quota/scaling limits, and integration failures
Design, develop, and deliver feature enhancements to existing applications
Support microservices architecture, APIs, authentication, inter-service communication, and failure/retry behavior
Own build, release, and deployment activities across development, test, pre-production, and production environments
Perform root-cause analysis and implement sustainable fixes
Build and maintain monitoring, alerting, logging, dashboards, and error reporting
Support platform, framework, library, dependency, and runtime upgrades
Support IAM, service accounts, access controls, secrets management, and operational governance
Participate in change management, release-readiness reviews, and on-call/rotational support
Collaborate with client and cross-functional engineering, product, QA, data, infrastructure, and platform teams
Create and maintain technical documentation, runbooks, troubleshooting guides, deployment procedures, and support playbooks
Requirements
3–7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related role
Strong hands-on experience supporting, maintaining, and enhancing production applications on Google Cloud Platform
Hands-on experience with Google App Engine deployment, configuration, scaling, versioning, and troubleshooting
Understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging
Experience with multi-environment deployments, release validation, rollback, and change control
Strong programming skills in one or more of Python, Java, Node.js/JavaScript, or Go
Working knowledge of SQL and application data stores including Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery
Understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations
Experience with CI/CD pipelines and automated build/deployment tooling such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI
Experience troubleshooting complex production environments and performing root-cause analysis under time pressure
Ability to understand existing systems, codebases, services, configurations, and client-specific tools and workflows
Strong communication and collaboration skills, including communication during user-impacting incidents
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.