The collaboration layer for AI teams: domain experts annotate and review agent behavior, engineers improve the agent from what they find.
-
Updated
Oct 3, 2026 - Python
The collaboration layer for AI teams: domain experts annotate and review agent behavior, engineers improve the agent from what they find.
Verifiable handwriting synthesis for failure-driven multimodal training and regression.
A powerful Python framework for writing and running portable regression tests and benchmarks for HPC systems.
Autonomous web browser agent that audits performance, functionality & UX for engineers and vibe-coding creators. 网站自主评估测试 Agent,支持 GUI/CLI 一键完成性能、功能使用与交互体验的测试评估
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Open-source physical observability for robotics workcells: deterministic replay, verifiable incident evidence, and generated regression tests.
A pytest plugin for automatically updating doctest outputs
Executable security regression testing for agentic applications and MCP-integrated systems.
Deterministic regression range for prompt-injection defences: probe classes, guard profiles, corpus verification, and diff mode.
千笔一文 Novel — 人 AI 共写的长篇网文创作台:条目化世界书 · 正则 must 契约(确定性复检,违规拦住锁定)· 三层审校(引证验真 + 投票)· 同人文外部导入可整批撤销 · 升级不动书稿 · BYOK 全本地 · MIT。A long-form fiction co-writing agent whose pipeline knows how to say no.
Drop-in pytest plugin for regression-testing AI agents — snapshot baselines, semantic comparison, mock LLMs
Your agent still returns the right answer -- but now it calls 3x the tools. Maida is the pre-merge behavioral regression gate for AI agents: baseline agent traces, gate PRs in CI, block behavior regressions before merge. Local-first, no cloud.
Useful decorators every Data Scientist should know
AI-driven Abaqus workbench: describe a model in a spec, get a real solved one back. Every capability is backed by a check you can run on your own seat, and the gates refuse a doubtful answer instead of returning it. AGPL-3.0 + commercial licence.
Record-and-replay for agent decision graphs: reproduce a prod agent failure as a committed regression test, and re-run your fix without live LLM calls.
DEPRECATED - TLS regression scanner for Firefox
Turn failed AI agent runs into replayable regression tests. Catch regressions before you ship.
🚀 Continuous testing report generator written in Python
8 character-driven AI coding agent skills for Codex: debugging, code review, regression testing, and release checks. Weird, but employed. Install: npx skills add SoonGwan/questionable-hires
CI regression gate for MCP servers: run a golden set of tool calls, diff against a baseline, fail the build when outputs get worse
To associate your repository with the regression-testing topic, visit your repo's landing page and select "manage topics."