latchhire

DevOps / Platform Engineer

Yassir · Algiers, Algeria (hybrid)
HybridNew mid devopsplatform engineer
Apply on Yassir →
Context & Vision In 2026, writing code is no longer the primary bottleneck; managing its complexity and ensuring its reliability is. We are building a highly resilient advertising platform with a very lean, senior internal team, and we intend to keep it that way. To achieve scale without the overhead of a large engineering department, we rely on an AI-first development paradigm and a philosophy borrowed from the best large-scale open-source projects. Our code is not public, but our governance is theirs: asynchronous communication, exhaustive written documentation, explicit rules, and uncompromising quality gates. It is the only way a small team of humans and AI agents ships serious systems without accumulating technical debt. This role owns the ground the whole platform runs on: the clusters, the pipelines, and the guarantees that everything ships predictably. The Role You will own the platform and delivery infrastructure end to end — the clusters, the infrastructure-as-code, the delivery pipelines, and the operational guarantees behind them. In a lean team, reliability is not a separate department; it is a discipline you carry for everyone. 1. Infrastructure as Code. Own the cloud footprint through infrastructure-as-code and a GitOps workflow. Infrastructure is declared, reviewed and versioned like any other code — no click-ops, no undocumented state. 2. Delivery Pipelines.Own CI/CD. Builds are reproducible, deployments are predictable, and rollbacks are boring. You make shipping a non-event. 3. Reliability, Performance & DR. Own observability (metrics, logs, distributed traces), performance testing, and backup / disaster-recovery. You define the SLOs that matter for a real-time serving platform and you make them measurable. 4. Quality Gates & AI-First CI. Enforce quality at every passage point in the pipeline. Integrating AI into CI/CD — automated reviews, security and policy checks — is an open frontier here, and we expect you to study it and propose implementations. Everything you build is documented; if it is not written down, it does not exist. The Tech Stack - Orchestration: managed Kubernetes on Google Cloud Platform. - IaC & GitOps: infrastructure-as-code and a GitOps workflow. - CI/CD: modern delivery pipelines (and proposals to evolve them). - Observability: metrics, logs, distributed tracing. - Runtime context: Rust services, a React frontend, streaming, relational databases. - Cloud:Google Cloud Platform. Profile & Requirements We are looking for a platform engineer who treats infrastructure as a product, owns reliability for the whole team, and is comfortable in a lean, high-quality, AI-first environment. This role is not suited to someone who wants to run a ticket queue inside a large ops team. Essential Experience - 4+ years in DevOps / platform / SRE roles, running production Kubernetes** for real workloads (GCP preferred). - Infrastructure-as-code and a GitOps workflow in production; infrastructure declared and reviewed as code. - Solid CI/CD ownership and a real observability practice (metrics, logs, distributed tracing) - Experience defining and defending SLOs, performance and disaster-recovery for latency-sensitive systems. Core Competencies & Mindset - AI Development Lifecycle: comfort integrating AI into the delivery workflow, and a point of view on automating quality and security gates in CI. - Uncompromising Reliability: a maniacal focus on predictability, recoverability and security. Surprises in production are the enemy. - Written & Asynchronous: exceptional written communication; runbooks and ADRs are part of the job, not a favour. Effective in a distributed, async environment (Paris timezone +/- 3h). - Ownership:** you carry reliability for the whole team and raise risks before they become incidents.
Posted 2026-09-10