Anahata ASI Benchmarks

The first benchmark is currently under development. Interested community members are encouraged to join the Anahata Discord group to propose prompts, discuss scoring dynamics, voting processes, or suggest sense of humour improvements.

The Industry's First Standardized Pure-Java AI Performance Leaderboard

Evaluating Large Language Models on real-world Java engineering: interactive Swing/JavaFX GUIs, LWJGL 3D graphics, AST refactoring, and JNA native C-library bindings.

🏆 Active Benchmark Suites

Explore the flagship benchmark programs and candidate model scorecards.

Benchmarks are community-driven and under continuous development. Anahata ASI users are encouraged to join the Anahata Discord to suggest new benchmark challenges.
Flagship Suite
Sponsor: NovaRouteAI Logo NovaRouteAI

Anahata-AGI-1

The master certification suite evaluating LLMs on multi-modal Java tasks, EDT context discipline, in-process safety, and JNA system hardware integration.

View Anahata-AGI-1 Results →

⚖️ Anahata ASI Judge Evaluations

Run scoring is determined exclusively by certified Anahata ASI judges.

Expert Consensus Scoring

Each benchmark execution is reviewed and scored on a 1–10 scale by Anahata ASI judges. The official benchmark scorecard reflects the arithmetic mean of all registered judge evaluations.

Qualitative & Architectural Review

Judges provide qualitative feedback on execution polish, code elegance, thread discipline, zero-defect autonomy, and creative presentation directly attached to the run telemetry.

Video Evidence & Verification

Every scored run is backed by full-resolution screen recording and YouTube video evidence, ensuring complete transparency and reproducibility across all evaluated models.