Senior Infrastructure Engineer leading infrastructure operations and stability for Buffer's social media tools. Focusing on modernization and developer tooling within a fully remote team.
Responsibilities
Keeping the lights on at a higher bar. You'll make our CI/CD pipelines faster and even more anti-fragile, ship boring deploys, and turn each incident into a new lesson rather than a repeat.
Modernizing the foundation. Buffer is not new, but our Infra is in continuous improvement. From deploying KEDA and Argo Rollouts to improving our signal-to-noise ratio in monitoring, there is plenty to do. One cannot talk about modernization without mentioning AI. We're already leaning on it for investigations and boilerplate, and want to bring it deeper into our day-to-day.
Treating engineers across Buffer as customers. You'll evolve the developer tooling that makes the inner loop fast, with documentation that holds up at 2am or when consumed by an agent.
Own the day-to-day reliability of our production platform. Keep EKS, ArgoCD, and the AWS surface area boring, tune autoscaling so the system adjusts well under load, and approach incident response in a way that each incident teaches us something new instead of repeating itself. (On-call is distributed across all engineers at Buffer, a week-long shift roughly once a quarter.)
Build progressive delivery into something the rest of engineering trusts. Implement Argo Rollouts with clean rollback paths, so the time between "this deploy is bad" and "this deploy is reverted" measures in seconds to minutes.
Build developer tools as products, instead of loose scripts. Evolve our in-house local development environment, BIBEs (Buffer Isolated Build Environments, per-PR full-stack staging deployments), and our CLI tooling so the inner loop is fast, frictionless, and parallel-friendly for AI agents. Measure adoption, talk to your users, iterate.
Reduce operational toil with AI. Automate low-risk workflows end-to-end so the team spends its time on the hard problems, not the repeat ones. AI doesn't touch infrastructure directly, it speeds up the humans who do.
Keep the stack current. Drive lifecycle upgrades: application runtimes (Node.js, Python), Kubernetes, EKS, Helm versions, and the Terraform-managed surface area. Own infra-side security vulnerabilities.
Improve the economics of our platform. Lead visibility work on Datadog, AWS rightsizing, and log filters so observability and cloud spend grow slower than the company does.
Partner with EPD on the platform they build on. Raise the documentation bar with the team, carry your share of weekly security work (dependency and vulnerability management is everyone's job), and help the infra team grow toward shared ownership and fewer single-person dependencies.
Requirements
You've worked as an Infrastructure Engineer, SRE, "DevOps" engineer, or adjacent role for long enough to be considered senior.
You have hands-on experience operating production Kubernetes at scale on a managed offering (GKE, EKS, AKS), including authoring and maintaining Helm charts, and you're fluent with autoscaling primitives driving KEDA and the cluster auto scaler.
You have AWS depth across IAM, EC2, S3, SQS, ECR, and ALBs. You may have also used Cloudflare (WAF, Workers, etc.) and GCP (BigQuery).
You have strong Terraform skills. You default to modules for structure, and keep the code adaptable, readable, and self-contained. Bonus points if you contributed an OSS module.
You've operated production CI/CD with GitHub Actions (or equivalent) and GitOps via ArgoCD (or similar). You've authored ArgoCD pipelines and Helm configuration yourself, including canary or progressive delivery systems you'd trust to roll back safely.
You've built internal developer tools (CLIs, dev environments, per-PR environments) and you think about them as products with users, not scripts.
You have a track record of pragmatic build-vs-buy decisions on infrastructure tooling. You can defend a choice and revisit it when conditions change.
You've worked with DataDog, Sentry, or similar observability stacks, and you design logs and metrics with cost in mind. You know observability and cloud spend can grow faster than the company if no one is watching.
You're comfortable with the Cloudflare across Workers, Zero Trust, DNS, and the rest of their platform.
You read and modify TypeScript or Node services well enough to upgrade runtimes and unblock teams (legacy PHP and Python show up too).
You're fluent with modern AI tools. You use them to debug, document, and reduce toil, not just to generate code, and you bring those patterns into how infra runs.
You're proactive and you follow through. You spot what needs doing before you're asked, and you close the loop without being chased.
You turn ambiguity into proofs of concept. You take fuzzy asks, ship something rough teammates can react to, and iterate with them until it lands.
You thrive in remote, asynchronous environments. You're clear in your thinking, generous with context.
You don't wait for perfect information to start, and you don't wait for perfect to ship.
You see infra as a force multiplier for engineering, not a gatekeeper.
You care about Buffer's customers. When things are slow for them it's painful for you to see. When errors are flaky you find the root cause and try to eliminate the entire class of problem, because you see the system, not the bug.
You care about performance. If it's too slow to use, it shouldn't exist. You'd rather make it fast than work around it.
You're a generalist engineer with strong spikes: T-shaped folks with depth in infrastructure and the flexibility to pivot as priorities shift.
You create, not just consume. Open source contributions, a technical blog, conference talks, side projects, or active accounts on the platforms Buffer serves, your pick. We're a Team of Creators ourselves, and the closer infra is to the creator's experience, the better the platform becomes.
You think about infrastructure as a platform with users. APIs, SDKs, CLIs, MCP servers, or developer-facing tooling you've shipped where adoption, not just deployment, was the success metric. You've felt the difference between code that ships and code that gets used.
You play the long game. You'd rather invest in compounding fundamentals than chase the platform-of-the-month.
Bonus points if you're already a Buffer user or familiar with social media management tools.
Senior Software Engineer building scalable Python and AWS infrastructure for d1g1t’s institutional wealth management platform. Driving implementation, deployment, troubleshooting, and large - dataset performance.
Senior Platform Engineer building resilient Go and Kubernetes platforms for Virtasant’s cloud - native products. Owning reliability, infrastructure as code, observability, and developer delivery tools.
Sr. Engineer position focusing on vulnerability remediation and cyber security for Genpact in Canada. Collaborating with teams and leading the cyber security team to meet business objectives.
Director of Infrastructure Engineering at Outschool leading infrastructure and platform team towards innovation and stability. Fostering a culture of technical excellence and ownership in a remote setup across U.S. and Canada.
IT Engineer responsible for IT infrastructure design and support in automotive software. Collaborating with teams to ensure efficient technology operations.
Senior Data Engineer designing and implementing scalable data lakehouse infrastructure at TRM Labs. Collaborating across teams to optimize data workflows and ETL processes.