An observability and evaluation platform for monitoring, testing, and optimizing generative AI and agentic systems.

Opik is an open-source platform for the observability, evaluation, and optimization of generative AI and agentic systems. It provides tools to track LLM calls, manage datasets, and monitor the performance of AI applications from prototype to production. The software is deployed as a server that can be self-hosted via Docker Compose for local development or Kubernetes for scalable production environments, or it can be accessed through a cloud instance.
Developers use the platform to log traces and spans to understand exactly how their models behave in real time. It includes a Prompt Playground for testing different model configurations and a set of SDKs for Python, TypeScript, and Ruby to integrate observability into existing codebases. The system is designed to handle high volumes of production data, with the capacity to support over 40 million traces per day for large scale deployments.
The platform is built for developers working with RAG chatbots, code assistants, and complex agentic workflows. It integrates with various LLM providers and orchestration frameworks to centralize data observability. By combining tracing with automated evaluation datasets, it allows teams to quantify the quality of AI responses and implement guardrails for responsible AI practices. The architecture supports various service profiles, allowing users to run only infrastructure or backend services depending on their specific development needs.
Opik serves as a technical infrastructure layer for LLMops, focusing on the full lifecycle of testing, monitoring, and refining AI systems.
A self-hosted control center for running autonomous coding agents to plan and ship changes across codebases.
A collaborative platform to build, schedule, and operate AI agents that handle long running automated tasks.
Join our newsletter to get shiny new open source software delivered to your inbox. Unsubscribe anytime.