Senior AI Engineer (Agents & Applications)
Firmus Technologies · Singapore
New
senior
ai engineer
Apply on Firmus Technologies →
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Why Firmus?
As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.
We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.
Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.
What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.
Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.
Role Summary
The Senior AI Engineer (Agents & Applications) will design, build, and operate production-grade agentic systems that coordinate, optimize, and automate decision-making across the design-build-operate lifecycle of AI factories. The role is a core contributor to the AI & Applications team’s Model-to-Grid product, connecting models, inference endpoints, benchmark intelligence, validated workload recipes, job-scheduler decisions, infrastructure telemetry, AI-factory operations, and grid-related constraints into safe, explainable, and measurable workflows.
The role will build more than conversational co-pilots. It will create agentic applications that ingest and reason over time-series telemetry, logs, traces, events, scheduler state, benchmark results, configuration data, operational documentation, incident records, and multimodal sources where appropriate. These applications will help engineers, operators, and customers move from observation to diagnosis, recommendation, planning, simulation, controlled execution, verification, and continuous improvement.
The engineer will define and implement the underlying agent architecture and engineering framework: orchestration, state and memory management, retrieval, tool use, specialized sub-agents, evaluation, safety controls, human approvals, observability, and deployment. The role will use fit-for-purpose self-hosted and external model endpoints, with close integration to the team’s inference platform.
Key Responsibilities
Design, build, and operate agentic applications supporting AI-factory planning, commissioning, validation, workload onboarding, benchmark analysis, model and recipe optimisation, scheduling, operations, maintenance, incident response, and continuous improvement.
Define reference architectures for single-agent, multi-agent, workflow-based, eventdriven, and human-in-the-loop agentic systems.
Build orchestration workflows using appropriate agent frameworks and libraries, such as LangGraph, LangChain, LlamaIndex, Microsoft AutoGen, Semantic Kernel, CrewAI, PydanticAI, Haystack, DSPy, or equivalent custom-built frameworks.
Select the appropriate architecture for each use case rather than applying multi-agent
patterns by default:
Deterministic workflow and state-machine architectures for repeatable, high-confidence operational processes.
Planner-executor architectures for decomposing complex investigation, planning, and remediation tasks.
Supervisor-worker or manager-worker architectures for coordinating specialist domain agents.
Router architectures for selecting the right model, tool, knowledge source, workflow, or specialist agent.
Reflection, critic, verifier, or judge patterns for quality assurance, validation, and safety checks.
Event-driven architectures for responding to telemetry anomalies, workload failures, scheduler events, benchmark regressions, and operational alerts.
Human-in-the-loop architectures for high-impact recommendations, privileged actions, or changes to production environments.
Build specialist agents for relevant Model-to-Grid and AI-factory domains, such as:
Benchmark-analysis and performance-diagnosis agents.
Workload recipe and runtime-configuration recommendation agents.
Inference-endpoint selection, capacity, and optimization agents.
Kubernetes and job-scheduler diagnostic agents.
GPU-topology, network, RDMA, storage, and utilization-analysis agents.
AI-factory health, capacity, maintenance, and operational-triage agents.
Documentation, knowledge, incident-review, and runbook-execution assistants.
Thermal domain specific monitoring and optimization agents.
Power domain specific monitoring and optimization agents.
Grid-integration specific monitoring and optimization agents.
Develop the intelligent coordination layer for Model-to-Grid, enabling agents to reason across model characteristics, inference and training configuration, validated recipes, GPU resources, topology, scheduling policies, network and storage performance, capacity, power, thermal conditions, health signals, and operational constraints.
Build Retrieval-Augmented Generation (RAG) pipelines using a combination of vector retrieval, hybrid search, metadata filtering, reranking, structured-data queries, graphbased retrieval where valuable, source attribution, and permission-aware access controls.
Use tools such as pgvector, OpenSearch, Elasticsearch, Milvus, Weaviate, Pinecone, Qdrant, Neo4j, or equivalent data and retrieval platforms as appropriate to the product architecture and deployment environment.
Design knowledge-ingestion pipelines for documentation, runbooks, ticketing systems, configuration repositories, benchmark reports, experiment records, cluster state, telemetry catalogues, incident reports, and approved internal knowledge sources.
Build data and context pipelines that combine unstructured knowledge with structured operational data, including metrics, logs, traces, events, time-series databases, scheduler queues, job states, resource inventories, and configuration-management data.
Integrate agents with governed tools and APIs, including Kubernetes, proprietary scheduler services, observability platforms, benchmark services, inference endpoints, configuration repositories, CI/CD pipelines, ticketing systems, workflow engines,
databases, and operational tooling.
Define tool contracts using structured input and output schemas, typed interfaces, validation, retries, idempotency controls, rate limits, timeouts, circuit breakers, approval requirements, and detailed audit logging.
Build robust agent harnesses that provide context assembly, model routing, prompt and policy versioning, structured output handling, memory management, state persistence, retries, failure handling, task recovery, escalation, and end-to-end tracing.
Implement short-term task memory, long-term user or operational memory where permitted, episodic memory for prior investigations or incidents, and semantic memory based on approved knowledge stores; apply retention, access-control, and datagovernance requirements to each.
Implement model-routing and fallback strategies across self-hosted inference endpoints and approved external models, selecting models according to task complexity, latency, cost, context-window requirement, tool-use capability, privacy needs, and reliability targets.
Partner with the Self-Hosted Inference Platform & Optimization team to ensure that agentic applications have suitable endpoint profiles for planning, reasoning, embeddings, reranking, summarization, classification, tool use, multimodal analysis, and high-throughput operational workflows.
Build agent workflows for detection, diagnosis, recommendation, planning, action simulation, controlled execution, verification, and learning loops.
Develop offline replay, simulation, shadow-mode, and what-if evaluation capabilities to validate recommendations before allowing actions in production—especially for scheduler policies, workload placement, runtime changes, capacity decisions, and operational remediation.
Design human approval and policy enforcement workflows that clearly present an agent’s evidence, recommendation, expected impact, proposed action, confidence, risk classification, authorization scope, and rollback option.
Work with the Security Engineer to implement defense-in-depth controls against direct and indirect prompt injection, insecure output handling, excessive agency, unsafe tool use, data leakage, cross-tenant exposure, privilege escalation, credential misuse, unauthorized actions, and insufficient auditability.
Use policy engines, guardrail frameworks, structured output validation, content and tool filters, permission checks, sandboxing, and allowlisted action patterns to ensure that agent behavior remains bounded and trustworthy.
Implement agent observability using tracing, metrics, logs, prompt and model version tracking, tool-call records, evaluation results, token and cost tracking, user feedback, incident evidence, and action audit trails.
Build and maintain evaluation frameworks using automated tests, curated test sets, simulation, replay, benchmark tasks, regression suites, model-based evaluators, human review, and operational-outcome measures.
Evaluate agents on task completion, factuality, groundedness, retrieval quality, diagnostic accuracy, recommendation quality, tool-selection accuracy, tool-execution correctness, policy compliance, latency, cost, safety, and user or operator satisfaction.
Partner with the Kubernetes and custom scheduler team to consume and explain queue state, placement rationale, topology information, capacity signals, workload lifecycle events, policy outcomes, and performance data.
Partner with the Model-to-Grid product, inference, Platform, SDI, Security, UX, and global operations teams to turn agent capabilities into clear product workflows, production releases, runbooks, and measurable user and operational outcomes.
Skills & Experience
5+ years of software engineering experience, including 3+ years building AI/ML applications, distributed systems, automation platforms, data products, or production workflow systems.
Demonstrated experience delivering LLM-powered, agentic, retrieval-augmented, decision-support, or operational-automation applications into production.
Strong Python expertise, including experience with FastAPI or comparable API frameworks, asynchronous programming, event-driven services, distributed task execution, data pipelines, and API integrations.
Hands-on experience with one or more agent frameworks, such as LangGraph, LangChain, LlamaIndex, Microsoft AutoGen, Semantic Kernel, CrewAI, PydanticAI, Haystack, DSPy, or an equivalent in-house agent framework.
Demonstrated ability to build custom agent orchestration when frameworks are insufficient, including state machines, graph-based execution, durable workflow execution, task queues, tool routers, model routers, planners, evaluators, and humanapproval flows.
Experience with workflow and orchestration technologies such as Temporal, Dagster, Prefect, Airflow, Argo Workflows, Kubernetes Jobs, Celery, or equivalent event-driven or durable-execution platforms.
Strong understanding of multi-agent architectures, including supervisor-worker, planner-executor, router, reflection, critic-verifier, debate, hierarchical, blackboard, and event-driven coordination patterns.
Practical experience with RAG architectures, including chunking, embeddings, vector search, hybrid retrieval, metadata filters, reranking, structured-data retrieval, SQL generation controls, knowledge graphs, provenance, and evaluation.
Experience with vector databases, search platforms, or graph databases such as pgvector, OpenSearch, Elasticsearch, Milvus, Weaviate, Pinecone, Qdrant, Neo4j, or equivalent technologies.
Familiarity with self-hosted and managed LLM inference, including model routing, OpenAI-compatible APIs, embeddings, reranking, tool calling, structured outputs, streaming, rate limits, latency, capacity, and cost management.
Understanding of model protocols and integration approaches such as REST, gRPC, WebSockets, event streams, OpenAPI, JSON Schema, and Model Context Protocol (MCP) or comparable tool-integration patterns.
Experience integrating AI systems with time-series data, logs, traces, observability platforms, ticketing systems, configuration repositories, CI/CD systems, databases, cloud services, and enterprise APIs.
Familiarity with cloud-native AI platforms, including Kubernetes, containers, workload scheduling, model serving, GPU resources, observability, multi-tenancy, and operational runbooks.
Understanding of agent security and responsible-AI controls, including prompt injection, indirect prompt injection, data and tenant isolation, tool authorisation, workload identity, permission boundaries, auditability, output validation, sandboxing, and humanin-the-loop safeguards.
Experience with agent evaluation and observability tools or patterns, such as OpenTelemetry, OpenInference, LangSmith, Langfuse, Arize Phoenix, Weave, TruLens, Ragas, DeepEval, promptfoo, custom test harnesses, or equivalent tooling.
Location & Reporting
Location: Singapore
Reports to: Head of AI & Applications
Employment Basis
Permanent full-time
Diversity
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionise the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.
Posted 2026-09-08