Observed arrival · 2026-09-26
FIBRE puts fixed-income models through their paces
FIBRE compares language models on fixed-income question answering and tool-assisted filing research, displaying scores alongside latency and estimated API cost.
- For
- Teams evaluating models for fixed-income analysis
- Worth noticing
- Its best-per-task comparison uses only the 200 tasks all five models attempted, and the page explains why that hindsight score cannot be reached by a router.
Field notes
The board contains 202 practitioner-written tasks: 152 brief-based question-answering items and 50 agentic-research items involving live filing research. Scores weight tasks equally; the leaderboard also separates track scores, median per-task wall-clock latency, and mean cost at listed API prices. Its “best per task” comparison takes the highest model score on each of 200 tasks attempted by all five models, and the page warns that this hindsight ceiling is unattainable by a router and rises with model count.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue