LibrariEval System
LeaderboardDocsPricingSign UpLogin
Librari Evals — Docs

Librari Evals — User Manual

Librari Evals is a system that allows you to develop, version, and benchmark structured-data extraction request schemas and their prompts across a comprehensive array of LLMs for any kind of document.

Use this system to:

  • Develop, maintain, version, and test structured-data request schemas on any document type you define (contracts, leases, ad infinitum).
  • Hit a single backend API for access to a wide array of models using the schemas you craft on our platform.
  • Build definitive benchmarks across a wide array of models — measure extraction results with your own ground truth documents, no more guessing about quality.
  • Monitor and be notified about LLM extraction quality over time, protecting against sudden quality drops by any provider or model (I'm talking to you, GPT-5.1!).
  • Measure the quality and impact of your extractions on any newly released models, immediately upon their release.

This manual covers the workflows used to run the system and use our API.

Start here

  • Concepts — definitions for every major term you'll see (Document Type, Request Schema Version, Ground Truth Response, Test Result Set, Rubric, etc.).
  • Workflow — how the pieces fit together end-to-end: build schemas → capture ground truth → benchmark → ship to production.
  • Build a Request Schema — the first concrete walkthrough; build your first schema