← All jobs

Agentic AI Engineer

Kivor Inc
Location
Remote / Hybrid / On-site (as per company policy)
Work mode
Remote
Employment
Full Time
Experience
3-5 years

About the role

This role focuses on designing and shipping production-ready, LLM-powered agentic systems. Key knowledge areas include agent orchestration (planning, tool calling, state management), integrating LLMs with external APIs and databases, building robust RAG and context pipelines, and ensuring agent reliability through comprehensive evaluation, guardrails, and observability. The position also emphasizes production optimization for latency, cost, and reliability using techniques like prompt engineering, caching, and selective fine-tuning.

Responsibilities

  • Build production agents that break down goals, plan, call tools/APIs, maintain memory, and complete multi-step workflows — with human-in-the-loop where it matters.
  • Design agent orchestration: tool/function calling, routing, planning loops, state management, and multi-agent collaboration.
  • Integrate LLMs with the real world — internal APIs, databases, third-party services, and MCP (Model Context Protocol) servers.
  • Make agents trustworthy: design evals for agent trajectories and outputs, add guardrails (prompt-injection defense, output validation, safe tool execution), and instrument tracing/observability.
  • Optimize for production: latency, cost, and reliability — via model selection/routing, prompt engineering, caching, and fine-tuning where justified.

Full description

Agentic AI Engineer Location: [City / Remote / Hybrid] · Type: Full-time · Level: Mid–Senior · Team: AI / Engineering About the role We’re looking for an Agentic AI Engineer to design and ship LLM-powered agents that plan, reason, use tools, and act autonomously to complete real multi-step tasks in production. You’ll own agentic systems end to end — from orchestration and tool integration to evaluation, guardrails, and reliability at scale. This is a hands-on build role for someone who has moved past demos and knows what it takes to make agents that are accurate, safe, fast, and cost-effective in the real world. You’ll work closely with product, design, and domain experts to turn ambiguous problems into dependable agentic products. What you’ll do Build production agents that break down goals, plan, call tools/APIs, maintain memory, and complete multi-step workflows — with human-in-the-loop where it matters. Design agent orchestration: tool/function calling, routing, planning loops, state management, and multi-agent collaboration. Integrate LLMs with the real world — internal APIs, databases, third-party services, and MCP (Model Context Protocol) servers. Build RAG and context pipelines: retrieval, chunking, embeddings, vector stores, and context-window management. Make agents trustworthy: design evals for agent trajectories and outputs, add guardrails (prompt-injection defense, output validation, safe tool execution), and instrument tracing/observability. Optimize for production: latency, cost, and reliability — via model selection/routing, prompt engineering, caching, and fine-tuning where justified. Ship and operate: deploy agents with CI/CD, monitoring, and incident response; iterate from real usage data. Collaborate: partner with product/design/domain experts and translate fuzzy requirements into robust agentic systems. What you’ll bring (required) 3+ years of software engineering experience, strong in Python (TypeScript a plus). Hands-on experience building and shipping LLM applications or agents in production — tool/function calling, RAG, and structured outputs. Familiarity with agent frameworks and patterns — e.g. LangGraph / LangChain, LlamaIndex, CrewAI, AutoGen, or custom orchestration — and with MCP. Experience with vector databases (pgvector, Pinecone, Weaviate, Qdrant, Chroma) and embedding models. Solid backend/API skills and cloud experience (AWS / GCP / Azure); comfort with async, queues, and distributed systems. Practical understanding of LLM behavior and limits — prompting, context windows, hallucination, evaluation — and how to engineer around them. Experience with evals and observability for LLM systems (offline + online), and a habit of measuring before shipping. Nice to have Multi-agent systems, planning/reasoning research, or reinforcement-learning exposure. Fine-tuning (LoRA/PEFT), model serving (vLLM/TGI), and working with open-weight models. Real-time / voice / streaming agents and low-latency inference. LLMOps / MLOps: data pipelines, prompt/version management, cost governance. Security mindset for agent safety: sandboxing, PII handling, prompt-injection and tool-abuse defense. Contributions to open-source AI tooling, or a portfolio of shipped agentic projects. Tech you’ll likely touch LLMs (Anthropic Claude, OpenAI, open models) · agent frameworks (LangGraph, etc.) · MCP · vector DBs · Python · FastAPI · cloud + containers · tracing/eval tools (LangSmith, Langfuse, Arize) · CI/CD. How we work / what we value Eval-driven development — you don’t ship an agent you can’t measure. Fast iteration and comfort with ambiguity and a moving field. Product sense — you care whether the agent actually helps the user, not just whether it runs. Ownership — from prototype to production to on-call. Optional sections to add: About the company · Compensation & equity · Benefits · Interview process · Equal-opportunity statement.
Not ready for a demo? Run a real role through a free pilot first.
Self-serve signup, no sales call, no charges during the pilot.
Start Free Pilot