Skip to content
View rudycelekli's full-sized avatar

Highlights

  • Pro

Block or report rudycelekli

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rudycelekli/README.md
Rudy Celekli — building AI systems that can prove what happened

LinkedIn Gradia ORCID Follow


Forward Deployed AI Researcher · Agentic AI Engineer

Building verifiable, long-horizon agent environments · Enterprise AI & GTM

Gradia · Snorkel AI · Axiom Consulting

Intelligence is cheap. Evidence is the product.

I’m a forward deployed AI researcher and agentic AI engineer building systems that make autonomous behavior observable, replayable, and independently verifiable.

My work sits where agent infrastructure meets experimental science: record what happened, challenge the evaluator, preserve the evidence, and make every important claim reproducible by someone else.

The proof loop: observe, record, replay, verify

Open-source impact, verified

My upstream work is discovered automatically across public repositories outside my account and included only when GitHub also lists me as a contributor. Merged PRs are the accepted-work measure; GitHub-indexed commits are a separate cached attribution signal. Stars and forks describe repository reach, not personal credit.

GitHub-verified contribution statistics for MoneyPrinterTurbo, openrig, agency-agents, Ruflo, Agentic-QE, CowAgent, VoiceStudio, iFixAi, OpenGTM, hindsight, benchmark-radar, imajev

Public projects I own

Owned public source is a separate signal: GitHub lists me as the repository owner. These projects are discovered automatically from my public, non-fork, non-archived repositories, then ranked by stars, recent activity, and GitHub-attributed commits. Ownership is not counted as upstream contributor credit.

Automatically discovered public projects owned by Rudy Celekli: ProofSeal, TestLore, Gradia Guard, ghostpart, nerve, gradia-wind-tunnel, gradia-universes-work-sample, gradia-reward-loop

  • ProofSeal: public owner repository with 13 GitHub-indexed commits. Repository reach: 2 stars and 0 forks; last pushed 2026-06-13.
  • TestLore: public owner repository with 99 GitHub-indexed commits. Repository reach: 1 star and 0 forks; last pushed 2026-10-01.
  • Gradia Guard: public owner repository with 10 GitHub-indexed commits. Repository reach: 0 stars and 0 forks; last pushed 2026-10-02.
  • ghostpart: public owner repository with 13 GitHub-indexed commits. Repository reach: 0 stars and 0 forks; last pushed 2026-10-01.
  • nerve: public owner repository with 3 GitHub-indexed commits. Repository reach: 0 stars and 0 forks; last pushed 2026-09-28.
  • gradia-wind-tunnel: public owner repository with 83 GitHub-indexed commits. Repository reach: 0 stars and 0 forks; last pushed 2026-09-03.
  • gradia-universes-work-sample: public owner repository with 34 GitHub-indexed commits. Repository reach: 0 stars and 0 forks; last pushed 2026-09-03.
  • gradia-reward-loop: public owner repository with 33 GitHub-indexed commits. Repository reach: 0 stars and 0 forks; last pushed 2026-09-03.

Showing 8 of 19 qualifying owned public repositories · browse every public source repository · complete evidence retained in machine-readable data

Last verified 2026-10-02 UTC · visual + evidence refreshed every 30 minutes by GitHub Actions · machine-readable evidence

The proof stack

🛡️ Gradia Guard

Proof-bound evidence for AI execution. Records hash-chained decisions and actions, verifies them offline, and makes the boundary of every claim explicit.

TypeScript Python AG-UI Cryptographic evidence

An evaluation immune system. Finds reward-hacking exploits, isolates their causal slice, and tests whether scorer patches truly converge.

Python Causal testing Offline replay DOI

An evidence-graded RL environment for multi-touch enterprise negotiation: 3,510 graded episodes, 13 models, and a fully reproducible methodology.

TypeScript Python RLVR Evaluation

Regression memory for coding agents. Seal behavioral claims once; check them over MCP or CI before a silent regression reaches history.

JavaScript MCP CI Deterministic testing

Agentic power, engineered

Agentic Power asks a practical question: how much accepted skilled work can one hour of human direction command?

I design for that multiplier, but I won’t publish one without exposing its scope, evidence, assumptions, and uncertainty. My operating model is to delegate meaningful responsibility, compose the right tools and context, verify the result, then use the evidence to improve the next run.

Live Agentic Power snapshot

Operator-calibrated provisional Agentic Power since 2025 with a month-by-month timeline

AP ≈ 78.3× (operator-calibrated provisional estimate): approximately 11,201.4 skilled Human-Equivalent Hours, or 280.0 engineer-weeks, divided by 143.0 operator-estimated human-direction hours. The transparent scenario range is 49.3× to 126.2×.

The evidence base since January 2025 combines 429 merged upstream PRs with 6,532 authored, non-merge default-branch commits across 109 currently accessible repositories: personal, private, employer, and open-source alike. Work such as Gradia is included only as a redacted aggregate: no repository names, commit messages, code, links, employer, or client details are published.

The conservative GitHub activity proxy produces 1,675.5 direction hours because it assigns attention to individual commits and repository-months. Operator recall is 1.5–2.0 active direction hours per week; across 81.7 weeks, that calibrates the denominator to 122.6–163.4 hours, with 143.0 as the midpoint.

Direction includes active briefing, steering, reviewing, correcting, and coordinating. It excludes agent runtime and waiting. The calibration is operator-estimated rather than reconstructed from time logs, so this remains a transparent scenario, not a completed Full Evidence Audit.

Framework and formula · calculation evidence · public upstream evidence and visuals refreshed every 30 minutes; redacted repository snapshot retained until a private read credential is available

Agentic operating model: delegate, orchestrate, verify, and improve

  • Direct: goals, constraints, acceptance criteria, and escalation rules stay explicit.
  • Orchestrate: models, specialist agents, reusable skills, tools, memory, and loops become one working system.
  • Verify: Gradia Guard, ProofSeal, and replayable evaluations separate completed work from plausible-looking activity.
  • Improve: Wind Tunnel and Reward Loop turn measured failure into a better next run.

Framework concept by Dr. Mark Allen / HeroForge.AI. Visual and operating-model adaptation are original to this profile.

What I’m exploring

Can we trust the record?      → proof-bound execution evidence
Can we trust the score?       → adversarial evaluator stress testing
Can we reproduce the result?  → deterministic replay + committed artifacts
Can an agent remember?        → regression memory before commit

The thread through all of it is simple: an AI system should be able to show its work without asking you to trust the AI system.

Working set

TypeScript Python Node.js MCP GitHub Actions Research

In the lab

  • Gradia Universes — proof-carrying, interruption-capable synthetic agent worlds with deterministic replay.
  • Gradia Reward Loop — oracle-witnessed reward-hacking experiments and a replay-verified paired-GRPO diagnostic.

Building the human layer

I build communities as deliberately as systems. I lead AI Tinkerers Charlotte, a demo-first room for people shipping real AI systems, and serve as Regional Ambassador for the Agentics Foundation, connecting builders working on practical agentic AI.


Building AI that can survive contact with reality.

If you care about agent reliability, evaluation integrity, or evidence-first AI, explore the work and compare notes.

Connect on LinkedIn →    Explore Gradia →    Research record →    Follow on GitHub →

Based in Charlotte · working in public

Popular repositories Loading

  1. proofseal proofseal Public

    Regression memory for coding agents — seal how your repo behaves, then check over MCP (or in CI) that an edit didn't regress it: pass/drift/regressed/missing. With seal history, bisection, and stal…

    JavaScript 2

  2. testlore testlore Public

    Build better tests. Run what matters. Remember what worked. An open-source test intelligence and quality engineering layer.

    JavaScript 1

  3. Vibe-Trading Vibe-Trading Public

    Forked from HKUDS/Vibe-Trading

    "Vibe-Trading: Your Personal Trading Agent"

    Python 1

  4. Papers-Literature-ML-DL-RL-AI Papers-Literature-ML-DL-RL-AI Public

    Forked from tirthajyoti/Papers-Literature-ML-DL-RL-AI

    Highly cited and useful papers related to machine learning, deep learning, AI, game theory, reinforcement learning

  5. data data Public

    Forked from jswansburg/data

    Data and code behind the articles and graphics at FiveThirtyEight

    Jupyter Notebook

  6. dr-streamlit-rudy dr-streamlit-rudy Public

    Forked from datarobot/dr-apps

    Python