Parallel achieves 70% accuracy on SEAL, benchmark for hard web research (parallel.ai)

🤖 AI Summary
Parallel announced that its Task API “Processor” family set a new price‑performance standard on SealQA (SEAL), a benchmark that stresses search‑augmented LLMs with conflicting, noisy, and ambiguous web research queries. Evaluated Oct 20–28, 2025 on SEAL-0 (111 questions) and SEAL‑HARD (254 questions) using an LLM-as-judge, Parallel reports 42%–70.1% accuracy across tiers and deterministic per‑query pricing (CPM). At the high end, Ultra8x reached 70.1% accuracy at $2,400 CPM (on SEAL‑HARD); the Pro tier hits 52.3% on SEAL‑0 and 66.9% on SEAL‑HARD at $100 CPM. These results outperform commercial baselines cited in the paper (e.g., Perplexity DR ~38–50%, Exa Research Pro ~45–59%, GPT‑5 ~49–65%) at multiple price points. Technically, Parallel credits systematic handling of web complexity: cross‑source conflict detection, credibility scoring that prefers primary/authoritative sources, high‑fanout exploration with disciplined pruning to manage compute, and an auditable Basis framework that returns citations, excerpts, structured reasoning, and calibrated confidences. The implication for AI/ML is twofold: (1) search‑augmented agents can materially improve factual synthesis under contradictory evidence when engineered for source reasoning and verification, and (2) consistent scaling across compute/price tiers suggests practical production readiness for use cases like due diligence, compliance, and competitive intelligence where defensibility and audit trails matter.
Loading comments...
loading comments...