Site Reliability Engineer responsible for Douro Labs' market data infrastructure performance and reliability. Develop automated workflows and lead incident responses for CeFi and DeFi users.
Responsibilities
Own the reliability and performance of Douro Labs’ real-time market data and price feed infrastructure powering CeFi and DeFi users
Monitor production systems, automate operations, and lead incident response with deep data forensics to keep services accurate and always on
Develop deep ownership of the end-to-end flow of a novel distributed system that generates and delivers real-time price feeds to hundreds of CeFi and DeFi users
Automate repetitive workflows, eliminate systematic issues, and proactively monitor live market data systems to prevent problems before they surface
Lead incident response and data forensics end to end, analyze impact, identify root causes, and ship tactical and long-term fixes that measurably improve reliability and resilience
Requirements
Strong ownership mindset with a passion for solving hard problems and operating mission-critical, always-on systems
Solid foundations in computer science, including data structures, algorithms, distributed systems, and time-series analysis
Proficient in at least one modern programming language (Python preferred) and comfortable working with SQL and production data
Able to reason about complex, adversarial markets and systems, identifying where risk and failure modes originate in real-time environments
Deep understanding of market data infrastructure, including ticker plants and reference data, with experience in asset pricing and liquidity analysis as a plus
Clear communicator who thrives in fast-paced, high-ownership environments and takes initiative to drive outcomes end-to-end
Benefits
Work remotely: Our team spans the globe from the US and South America to Europe and Asia
English proficiency is essential, as it's our primary language for team and external developer communication
Thrive in a startup environment within the dynamic DeFi space
Write open-source software: Check out our open-source projects on GitHub for a sneak peek into our work
User-centric development: Our success hinges on external developers engaging with our solutions
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.