Skip to content
SHIFTby Guruge
All work

AI / Developer Tool 2026

LLM Reliability Lab

A local-first lab for running models and prompts against the same datasets, and evaluating which configuration actually did better.

Year
2026
Status
In development
Platform
Web · API · Local models
Scope
Web application, API, Evaluation engine
reliability-lab / pipeline
  1. 01DatasetItems imported and validated
  2. 02ExperimentModel, prompt template, parameters
  3. 03RunBounded concurrency, streamed progress
  4. 04EvaluationPluggable evaluators
  5. 05ComparisonMetrics and regression detection
  6. 06RAG and tracingPlanned

Overview

LLM Reliability Lab is an engineering tool for developers building on language models. It records datasets, experiments and runs, then scores them — so a comparison between two models or two prompts is reproducible rather than a feeling.

Challenge

Comparing LLM configurations is usually ad hoc: a notebook here, a spreadsheet there, and no shared record of what was tried or why one version beat another.

Direction

Treat it as a developer tool, not a chatbot. Application code depends on a provider interface rather than on any one model API. Experiments are fixed configurations, runs are recorded per item, and evaluation is pluggable and kept separate from generation.

What was built

  • A provider-agnostic LLM interface, implemented for local models through Ollama
  • Model discovery with distinct loading, empty and unreachable states
  • Playground with streamed output, or structured JSON validated against a schema
  • Datasets with JSON and JSONL import and per-line validation errors
  • Experiments run across a dataset with bounded concurrency, live progress and cancellation
  • Side-by-side run comparison with a word-level diff
  • Evaluators: exact match, contains, local-embedding similarity and LLM-as-judge
  • Aggregate metrics and baseline-versus-candidate regression detection

Technology

  • Next.js
  • TypeScript
  • Tailwind CSS v4
  • FastAPI
  • Pydantic v2
  • SQLAlchemy (async)
  • PostgreSQL
  • Ollama

Outcome

Four of ten planned phases are complete: foundation, execution, experiments and evaluation. A model or prompt change can now be run across a dataset and checked for regressions against a baseline.

Technical notes

  • Local by default

    Models run through a local Ollama instance and semantic similarity uses a local embedding model, so evaluation needs no paid API.

  • Not yet built

    Retrieval (RAG), tracing and model routing are on the roadmap. The RAG package exists as a placeholder and is not presented as a feature.

Start a project

Want something like this?

Tell me about the business and what the site needs to do.