Senior SRE Engineer
Flowcode · New York, Hybrid
HybridNew
senior
sre engineer
Apply on Flowcode →
Senior SRE Reliability Engineer
Location: New York, NY (Hybrid) / Remote
Department: Engineering
The Role
Flowcode is seeking a Senior Site Reliability Engineer (SRE) to work on reliability and infrastructure efforts across our platforms. This role will help grow and drive our infrastructure strategy, operational rigor and observability while building and supporting the systems and tooling required to support Flowcode’s continued growth.
As an individual contributor within our engineering organization, you will develop and operate scalable cloud infrastructure, establish best practices around deployment and reliability, and partner closely with engineering teams to ensure systems are scalable, resilient and observable.
What You’ll Do
Reliability & Infrastructure
Improve system availability, scalability, and resilience across Flowcode's platforms
Own key pieces of our EKS-based infrastructure end-to-end
Contribute to incident response and postmortems, turning findings into durable fixes
Support engineering teams with infrastructure questions, escalations, and day-to-day unblocking
Cloud & Platform Engineering
Manage and scale our core AWS footprint (EKS, VPC, RDS) through Infrastructure as Code (Terraform)
Enhance disaster recovery and failover mechanisms to protect mission-critical workloads
Collaborate with product engineering to streamline and optimize internal developer experience
CI/CD & Deployment Automation
Design and scale deployment pipelines using GitHub Actions
Expand GitOps practices and tooling through ArgoCD
Facilitate secure delivery with automated validation and progressive rollout strategies
Observability & Monitoring
Oversee and optimize the organization's monitoring, logging, and alerting infrastructure
Develop high-signal metrics, tracing, and visualization dashboards while minimizing operational noise
Establish and monitor Service Level Objectives for managed platform components
Qualifications
Required
4+ years of professional experience across SRE, DevOps, or Platform Engineering domains
Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs
Advanced Terraform or OpenTofu expertise, encompassing module architecture and production state management
Hands-on operational experience with GitOps workflows via ArgoCD and Helm-based deployments
Ability to author production-grade code in Go or Python alongside robust shell scripting
Mastery of core AWS services, specifically EKS, Networking/VPC, RDS, and IAM
Experience maintaining and scaling CI/CD automation using GitHub Actions within collaborative environments
Proven track record of leading infrastructure initiatives from initial design through to long-term operation
Background in supporting large-scale distributed systems within high-availability production environments
Adept at navigating interrupt-driven workflows, balancing strategic project delivery with day-to-day operational support
Preferred
Exposure to Crossplane or alternative Kubernetes-native solutions for infrastructure provisioning
Deep observability experience utilizing Datadog or Prometheus to engineer SLOs, high-signal dashboards, and intelligent alerting
Practical knowledge of modern secrets management frameworks and implementation
Experience optimizing cluster efficiency through autoscaling technologies such as Karpenter or Cluster Autoscaler
Flowcode is not for everyone. We hire with a pinhole lens — only those with the rare combination of intellectual horsepower, execution velocity, and uncompromising drive will thrive here. If you are seeking to operate at the highest levels of performance and impact, we want to meet you.
How to Apply
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
A successful candidate’s starting pay will be determined based on the role, job-related skills, experience, qualifications, work location, and market conditions.
Posted 2026-08-25