Millions of production systems still run on legacy PHP: banks, insurers, government services, e-commerce platforms built 15+ years ago. That code works, but it's stranded: modern data science, machine learning, and AI tooling lives overwhelmingly in the Python ecosystem. Rewriting by hand is slow, expensive, and risky. Every subtle behavioral difference is a potential production incident.
php2python explores an automated path: translate a multi-file PHP project into Python using LLM agents, then prove the translation behaves identically by running both programs and comparing their output, not by trusting the model.
The core problem it tackles is cross-file consistency. Translating one file is easy; translating five files that call each other is where LLMs fall apart: renamed methods, dropped arguments, invented code. This pipeline treats the model as an unreliable worker inside a system of deterministic checks, rather than as a trusted translator.
For each PHP file, main.py orchestrates a chain of gpt-4o-mini agents:
- Symbol Table Agent: reads every PHP file and builds a JSON contract of the whole project: every class, method, attribute, and relationship (inheritance, interfaces, traits). This becomes the single source of truth all later agents must obey.
- Transformer Agent: translates the file to Python, required to match the symbol table exactly (names are immutable; other files depend on them).
- Verifier Agent: statically inspects the generated Python for undefined symbols, broken relationships, and type inconsistencies, returning structured JSON issues.
- Repair Agent: fixes only the reported issues, minimally.
Translated class files become live Python modules (generated/*.py,
registered in sys.modules so imports between them resolve). The entry
point runs with its output captured, then the original PHP runs via the
php CLI and the two outputs are compared after normalization
(timestamps masked, numeric formats normalized: PHP's 100 equals
Python's 100.0).
The finish line is literal: Outputs are equal!
Prompt rules alone don't make a small LLM reliable. This repo's answer is
a set of deterministic mechanisms in main.py that catch and correct the
model's recurring failure modes in code, where compliance isn't optional:
- Class ownership & dedupe (
build_class_ownership,remove_duplicate_classes) : the model loves inlining stub copies of other files' classes (a cut-downOrderinsideCustomer.py), which breaksisinstancechecks and drops methods. Ownership is derived from the PHP sources themselves; any class defined outside its owning module is deleted and replaced with a real import. - Fabrication stripping (
strip_fabricated_attr_resets) : the entry-point PHP only ever calls methods, so any generatedobject.attribute = ...statement in Main is a hallucination that clobbers constructor state. Matching lines are removed mechanically. - Bad-import removal (
strip_bad_import) : imports of nonexistent modules (from Logger import Logger) are detected from Python's ownModuleNotFoundErrorand deleted, no LLM call needed. - Traceback-targeted self-healing : when generated code crashes, the
pipeline walks the traceback to find every generated file involved
(the root cause may be the caller or the callee), feeds the real error
to a repair agent, rewrites the culprit files, and reloads all modules
in dependency order so fixes propagate to their importers
(
culprit_generated_files,module_file_for_missing_attr,reload_all_modules). - Output normalization (
normalize) : masks wall-clock timestamps and normalizes values on both sides so the comparison tests behavior, not incidental formatting, while preserving line structure, so real differences still fail.
Every generated file is written to generated/ before execution and
compiled with its real path, so crash tracebacks point at actual files
and lines instead of <string>.
- Python 3.9+ with the
openaipackage (pip install openai) - PHP CLI on PATH (
php -vto check): used to produce the reference output OPENAI_API_KEYset in the environment
export OPENAI_API_KEY="your-key"
python3 main.py