Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Responsibilities
Provide production support and maintenance for enterprise applications hosted on Google App Engine
Triage, diagnose, restore service, and drive user-impacting incidents to closure within agreed SLAs/SLOs
Troubleshoot application errors, latency, performance degradation, service failures, configuration issues, quota/scaling limits, and integration failures
Design, develop, and deliver feature enhancements to existing applications
Support microservices architecture, APIs, authentication, inter-service communication, and failure/retry behavior
Own build, release, and deployment activities across development, test, pre-production, and production environments
Perform root-cause analysis and implement sustainable fixes
Build and maintain monitoring, alerting, logging, dashboards, and error reporting
Support platform, framework, library, dependency, and runtime upgrades
Support IAM, service accounts, access controls, secrets management, and operational governance
Participate in change management, release-readiness reviews, and on-call/rotational support
Collaborate with client and cross-functional engineering, product, QA, data, infrastructure, and platform teams
Create and maintain technical documentation, runbooks, troubleshooting guides, deployment procedures, and support playbooks
Requirements
3–7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related role
Strong hands-on experience supporting, maintaining, and enhancing production applications on Google Cloud Platform
Hands-on experience with Google App Engine deployment, configuration, scaling, versioning, and troubleshooting
Understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging
Experience with multi-environment deployments, release validation, rollback, and change control
Strong programming skills in one or more of Python, Java, Node.js/JavaScript, or Go
Working knowledge of SQL and application data stores including Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery
Understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations
Experience with CI/CD pipelines and automated build/deployment tooling such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI
Experience troubleshooting complex production environments and performing root-cause analysis under time pressure
Ability to understand existing systems, codebases, services, configurations, and client-specific tools and workflows
Strong communication and collaboration skills, including communication during user-impacting incidents
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.