
Spanlens
Langfuse
Helicone AI
LangSmith
Comet.com
Evidently AI
Langtrace AI
Humanloop
Langfuse
Hugging Face
LangSmith
Helicone AI
LangChain
LastMile AI
ChatGPT
Spanlens is an open source observability tool for LLM apps. You point your OpenAI, Anthropic, or Gemini client at the Spanlens proxy by changing the baseURL, and it records every request with the full body, token counts, cost, and latency.
The dashboard shows per-model costs, latency percentiles, and error rates. Agent workflows appear as traces with a timeline view and a graph view that marks the critical path. There is also prompt versioning with A/B experiments, an LLM-as-judge eval runner, anomaly alerts, and a scanner that flags PII and prompt injection in request bodies.
The whole codebase is MIT licensed. You can use the hosted version or run it yourself with Docker Compose. SDKs exist for JavaScript and Python, and OpenTelemetry traces can be ingested over OTLP.
Spanlens
HumanloopNo Spanlens videos yet. You could help us improve this page by suggesting one.
Spanlens's answer
I was building LLM apps on the side and kept pasting token counts into a spreadsheet to figure out what each feature cost me. The tools I tried were either acquired mid-migration, closed source, or heavier to self-host than the app I was trying to monitor. So in April 2026 I started building the tool I actually wanted: change one line, see every request. It launched in June 2026, and the whole codebase went up on GitHub under MIT from day one.
Spanlens's answer
The entire product is MIT licensed, including the dashboard, evals, and prompt A/B testing. There is no separate enterprise edition. Everything ships in one repo you can run with a single Docker Compose file. Integration is one line: you change the baseURL on your OpenAI, Anthropic, or Gemini client, and every call gets logged with its full body, token counts, cost, and latency. A few things that are usually paid add-ons come built in, like agent traces with a critical path view, A/B tests that use Welch's t-test to tell you whether a difference is real, and a recommender that flags cheaper models based on the traffic you actually send.
Spanlens's answer
Mostly because of where the market went. Helicone was acquired, LangSmith is closed source, and self-hosting Langfuse takes real setup work. Spanlens fills the gap those tools left: setup in about five minutes, one Docker Compose file if you want the data on your own servers, and no feature gating between free and paid tiers. To be fair, if you need SOC 2 reports and enterprise support today, the bigger platforms are ahead. If you want request logs, costs, and traces without adopting a heavy platform, that is what Spanlens was built for.
Spanlens's answer
Developers who ship LLM features in production apps. The typical user is a solo developer or a small team that added OpenAI or Anthropic calls to their product and now has no clear picture of what those calls cost or why some are slow. Agent builders are the second group, since multi-step workflows are hard to debug without traces. It is a developer tool through and through: if you don't touch code, you won't get much out of it.
Spanlens's answer
We don't publish customer names yet. Spanlens launched in June 2026, and most users so far are indie developers and small AI teams.
Spanlens's answer
TypeScript across the stack. The dashboard is Next.js, the API and proxy run on Hono, and data is split between Supabase Postgres for accounts and relational data and ClickHouse for request logs, which grow fast. The repo is a pnpm monorepo that also holds the JavaScript and Python SDKs, a CLI, and an MCP server. Self-hosting runs on Docker Compose, and OpenTelemetry traces can be ingested over OTLP.
Based on our record, Humanloop seems to be more popular. It has been mentiond 5 times since March 2021. We are tracking product recommendations and mentions on various public social media platforms and blogs. They can help you identify which product is more popular and what people think of it.
Humanloop | London and San Francisco | Full time in person | https://humanloop.com Humanloop is building infrastructure for AI application development. We're the LLM Evals Platform for Enterprises. Duolingo, Gusto, and Vanta use Humanloop to evaluate, monitor, and improve their AI systems. ROLES:. - Source: Hacker News / over 1 year ago
- https://humanloop.com/) for teaching me the philosophy of implementing a copilot textarea. I wish I could have used the project directly, but integrating just one React component into Rails while keeping importmap and StimulusJS was quite challenging. Given the limited time, I decided to move on with StimulusJS. This is our first time building an open-source project to share with the world, and weโre a bit... - Source: Hacker News / almost 2 years ago
- Conversational simulation is an emerging idea building on top of model-graded evalโ - AI Startup Founder Things to consider when comparing options: โTypes of metrics supported (only NLP metrics, model-graded evals, or both), level of customizability; supports component eval (i.e. Single prompts) or pipeline evals (i.e. Testing the entire pipeline, all the way from retrieval to post-processing)โ โ+method of... - Source: Hacker News / almost 3 years ago
Humanloop (YC S20) | London (or remote) | https://humanloop.com We're looking for exceptional engineers that can work at varying levels of the stack (frontend, backend, infra), who are customer obsessed and thoughtful about product (we think you have to be -- our customers are "living in the future" and we're building what's needed). Our stack is primarily Typescript, Python, GPT-3. Please apply at... - Source: Hacker News / over 3 years ago
https://humanloop.com/ Find the prompts users love and fine-tune custom models for higher performance at lower cost. - Source: Hacker News / over 3 years ago
Langfuse - Langfuse is an open-source LLM engineering platform that helps teams collaboratively debug, analyze, and iterate on their LLM applications.
Helicone AI - Open-source LLM Observability for Developers
Hugging Face - The AI community building the future. The platform where the machine learning community collaborates on models, datasets, and applications.
LangSmith - Build and deploy LLM applications with confidence
Comet.com - Build better models faster
Evidently AI - Open-source monitoring for machine learning models