|
Bandito — Autonomous AI Engineer
Bandito connects to your LLM traces and turns them into improvements. The trace logging space is commoditizing — Bandito is the intelligence layer that sits...
|
80 % |
|
EvalDog — catch LLM regressions before your users do
Hosted prompt & RAG evaluation. Upload your test cases, get a graded pass/fail report in seconds. Plus a zero-token CLI for CI and AI agents.
|
80 % |
|
Cailos · The Decision Engine for Agentic AI
Hot-swap models across 20 providers and 107 endpoints for best-in-class task-tuned performance. Optimize for quality, speed, or cost.
|
79 % |
|
modelux — the control plane for LLMs
Policy-driven routing, finance-grade budgets, full decision traces, replay — across every LLM provider, without changing your code.
|
79 % |
|
Crane | Token Relay
Route 30+ frontier models through a single OpenAI-compatible endpoint. Millisecond routing, automatic failover, token-based billing.
|
79 % |
|
errorbar – LLM quality, with error bars
errorbar measures the quality of your LLM application on your own traffic and acts on it: judges calibrated against your grades, non-inferiority gates on can...
|
79 % |
|
Microscope — Stop Chargebacks Before They Happen
AI-driven chargeback prevention for Shopify and DTC merchants. A four-stage risk engine — Radar, Sniper, Judge, Report — with an order-level evidence chain b...
|
79 % |
|
Model Olympics
Which AI model wins your prompt? One prompt, three models, real metrics.
|
78 % |
|
Alignmenter: Test Your AI's Voice and Safety
Open-source testing tool for AI behavior. Check if your AI matches your brand voice, stays safe, and behaves consistently across updates.
|
78 % |
|
Model Maverick — the LLM model recommendation engine
Model Maverick picks the right LLM for every prompt — benchmarked nightly against a private task corpus, ranked under your policy, and auditable end to end.
|
78 % |
|
Provara — The Adaptive LLM Gateway
Routes every request. Learns from every response. Catch provider regressions, cut spend at equal quality, and answer 'why did our bill double?' in one screen...
|
78 % |
|
Vynaris — same quality, a fraction of the cost
Cut AI costs drastically without compromising on privacy or reliability. Vynaris routes each request to the cheapest model that clears your quality bar.
|
78 % |
|
AI Workflow Builder: Build, A/B Test and Deploy | Evaligo
Visual drag-and-drop AI workflow builder with built-in A/B testing. Run a flow on several AI models, score every output with blind AI judges, train a per-inp...
|
78 % |
|
Data2Distance: Turn Launch Data Into Lower Scores
Brand-agnostic golf equipment optimization from your own launch-monitor numbers. Find your driver's power leak free in 60 seconds, then get the exact specs t...
|
78 % |
|
Perssona · Orchestrate multiple LLMs in your terminal
Multi-agent. Multi-provider. A single canvas. Claude, GPT, Gemini and more — pick the brain, split the work, pay much less. For devs who want speed.
|
78 % |
|
ModelWatch — Catch silent LLM drift before your customers do
Continuous behavioral drift monitoring for LLM-powered apps. Catches provider-side model updates that quietly break GPT-4 / Claude in production. Bring your...
|
78 % |
|
AutoAW — Auto Agentic Workflows
AutoAW automatically discovers optimal multi-agent workflow configurations — co-evolving topology, prompts, models, and tools to find the cheapest pipeline t...
|
78 % |
|
ModelGate — cost control for LLM APIs
An LLM gateway for OpenAI, Anthropic, Google and Azure. Every request logged, priced to the token, and audited for waste you can actually recover.
|
78 % |
|
Farka | Fine-tune specialized AI models on your own data
Fine-tune small, specialized AI models on your own data. Every run is scored against the base model before export, and the weights are yours.
|
78 % |
|
SystemLore — Distributed systems simulation game
Learn distributed systems fundamentals through a simulation game: keep services alive through traffic spikes, outages, queues, caches, shards, and failover d...
|
77 % |
Detected social media networks