Skip to content
#

code-benchmark

Here is 1 public repository matching this topic...

Deterministic evaluation and reproducible benchmarking for AI-generated code with hardened Docker sandboxing, trusted tests, static analysis, and multi-provider model comparison.

  • Updated Sep 1, 2026
  • Python

Add this topic to your repo

To associate your repository with the code-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more