Ansh Roshan
Back to projects

Gen AI · 2026

LLM Router

A cache-aware LLM router with learned quality routing: one decision engine, two runtimes (pip + npm), zero dependencies, and an HTTP /decide sidecar that slots into LiteLLM and Bifrost to pick the best model per request on cost, latency, and quality.

PythonTypeScriptLiteLLMBifrost
LLM Router

More projects

View all
Loupe — AI PR Reviewer
Gen AI

A self-built AI pull-request review agent in the spirit of CodeRabbit: a zero-dependency review engine that reads any diff, posts inline comments plus a summary, and runs as a GitHub Action or Cloudflare Worker with your own LLM key — no server required. Named for the jeweler's lens: close, careful inspection of every diff.

TypeScriptGitHub ActionsCloudflare Workers+1
OmniBench — LLM Benchmark
Gen AI

Point it at any model — API, aggregator, or local — and it runs a 5-pillar battery (capability, reliability, safety, agency, economics), recording every run, task, turn, span, and token. Compare models side-by-side in a full dashboard with CLI/TUI runners, three.js 3D data views, and portable run bundles that pool into community averages.

TypeScriptthree.jsCLI/TUI+1
RAGStack
Gen AI

A hybrid agentic RAG system: it reads PDFs, Office docs, markdown, code, and web pages, builds four kinds of search indexes — BM25 vectorless, LanceDB vector, an LLM-extracted knowledge graph, and read-only text-to-SQL — and an agent decides which to chain for each question, always answering with citations. Local-first, MIT.

PythonBM25LanceDB+2