Show HN: We built the first comprehensive benchmark for legal retrieval (huggingface.co)

🤖 AI Summary
Isaacus today released the Massive Legal Embedding Benchmark (MLEB), a 10-dataset suite designed to be the largest, highest-quality benchmark for legal text embeddings. MLEB covers diverse jurisdictions (US, UK, Australia, Ireland, Singapore, EU), document types (decisions, legislation, regulations, contracts, literature) and tasks (retrieval, zero-shot classification, QA). Seven of the ten sets are newly created or expert-labeled; notable additions include the Australian Tax Guidance Retrieval dataset, which pairs 112 real taxpayer forum questions with government guidance, giving the benchmark realistic, hard user queries. The authors explicitly designed MLEB to address shortcomings in prior collections—limited coverage in LegalBench‑RAG and mislabeling/diversity problems in MTEB’s legal split—and released permissively licensed data, evaluation code, and raw results for reproducible research. Benchmarking on MLEB highlights the value of legal domain adaptation. Kanon 2 Embedder tops the leaderboard with NDCG@10 = 86%, slightly ahead of Voyage 3 Large (85.7%), while also posting the fastest inference times among tested commercial models (≈4× faster than Voyage 3 Large), claiming a new accuracy–latency Pareto frontier through parameter efficiency and licensed legal training data. Results show that strong MTEB performance doesn’t guarantee legal retrieval strength (e.g., Gemini ranks highly on MTEB but lower on MLEB), and warn of dataset contamination risks when models are trained on customers’ private data. MLEB thus provides a more realistic, reproducible yardstick for building and evaluating legal embeddings and highlights the practical gains from specialized, high-quality domain adaptation.
Loading comments...
loading comments...