Research

Benchmarks and papers on evaluating AI systems against the claims they actually make.

ScopeBench-PR

★ Best Paper Presentation

Measuring whether LLM reviewers respect a paper’s stated scope.

GlobalSouthAI · IJCAI–ECAI 2026

ICFD-31k

A large-scale benchmark for real-time conversational fraud detection.

IJCAI–ECAI 2026 · AI for Social Good Special Track

Temporal Retrieval

Selective retrieval for time-series forecasting beyond long-context scaling.

ICLR 2026 · TSALM Workshop