Skip to content
#

eval

Here are 212 public repositories matching this topic...

benchmark-radar

Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.

  • Updated Oct 5, 2026
  • Python

A self-evolving Agent OS — every interaction sharpens how it judges, not just what it knows. Four ideas: self-evolution (mistakes become code gates, not logged lessons), brain-first domain brains that grow and decay, cognition that persists between sessions, and proprioception — it senses and drives its own body. Human directs. AI delivers.

  • Updated Oct 5, 2026
  • Python

Add this topic to your repo

To associate your repository with the eval topic, visit your repo's landing page and select "manage topics."

Learn more