latchhire

Staff Site Reliability Engineer

Caseware · Toronto, ON (remote)
RemoteNew staff site reliabilityreliability engineer
Apply on Caseware →
Caseware is one of Canada's original Fintech companies, having led the global audit and accounting software industry for over 30 years, with more than 500,000 users across 130 countries and available in 16 different languages. While you might not have heard of us (yet) over 36,000 accounting and audit professionals list Caseware as a skill on their LinkedIn profiles! This is a hands-on senior engineering role focused on improving production resilience, strengthening security, driving operational excellence, and enhancing the developer experience across the organization. In this role, you will design, build, and evolve the foundational systems, tooling, and operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence. You will help establish reliability standards, define service level objectives (SLOs), improve observability, automate operational processes, and drive incident management and post-incident learning practices that strengthen platform stability over time. Partnering closely with Engineering, Security, Platform, and Product teams, you will architect scalable distributed systems, optimize Kubernetes and AWS-based infrastructure, and build automated delivery pipelines that support rapid and safe software releases. You will play a key role in reducing operational toil, improving system performance, increasing platform reliability, and ensuring that our infrastructure can support continued business growth. ❗ This is a full-time permanent position ❗ This is an existing vacancy 📍 Location: This is a remote location open to candidates legally authorized to work in Canada.
Posted 2026-08-18