AI Code Generation Benchmarks
Empirical comparison evaluating large language models generating orchestration code across Autolang (Strict and Loose modes), JavaScript, and Python. The results evaluate first-pass success rate, error categorization, retry behavior, and runtime failure rates.
Figures are specific to this run. First-pass rates depend on the model as much as on the language — treat them as one measurement point, not a standing guarantee.
0.25% Runtime Failures
Across all tested models and 400 runs per environment, Autolang recorded only 1 runtime error in each mode (0.25%), compared to 10 runtime crashes in JavaScript.
95.3%
In loose mode, Autolang achieved 95.3% accuracy after retry (381/400), outperforming JavaScript (93.8%) and exceeding its initial pass rate with targeted corrections.
94.0% Pass@1
Autolang loose mode reached a 94.0% zero-shot first-pass success rate (376/400), higher than JavaScript (92.5%) and approaching Python (95.0%).
Overall Environment Comparison
Aggregated results across all test suites, evaluating overall reliability, error distribution, token usage, and execution latency.
| Environment | Pass@1 | Pass After Retry | Avg Retry | Nominal Rate | Logic Err | Compile Err | Runtime Err | Infra Err | Avg Tokens | Avg Latency |
|---|---|---|---|---|---|---|---|---|---|---|
| autolang-strict | 88.5% (354/400) | 94.3% (377/400) | 1.87 | 94.3% (377/400) | 12 | 10 | 1 | 0 | 270 | 19.9 ms |
| autolang-loose | 94.0% (376/400) | 95.3% (381/400) | 4.75 | 95.3% (381/400) | 16 | 2 | 1 | 0 | 270 | 19.8 ms |
| javascript | 92.5% (370/400) | 93.8% (375/400) | 2.60 | 93.8% (375/400) | 10 | 5 | 10 | 0 | 270 | 1.1 ms |
| python | 95.0% (380/400) | 97.0% (388/400) | 2.38 | 97.0% (388/400) | 10 | 0 | 2 | 0 | 270 | 38.3 ms |
Runtime Reliability: Unlike JavaScript which recorded 10 uncaught runtime crashes, Autolang shifts verification to deterministic compilation and capability bounds. Scripts either compile cleanly against granted host capabilities or fail deterministically at compile time, yielding only 1 runtime failure across 400 tasks.
Strict vs Loose Modes: autolang-loose relaxes strict implicit variable and bracket closing constraints, boosting Pass@1 from 88.5% to 94.0% and total success rate after retry to 95.3%.
Detailed Model and Environment Results
Granular performance across diverse open-weights and commercial models, measuring code generation fidelity and prompt efficiency.
| Model | Environment | Pass@1 | Pass After Retry | Avg Retry | Nominal Rate | Logic | Compile | Runtime | Infra | Prompt Tok | Resp Tok | Total Tok |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gpt-oss-20b | autolang-strict | 98.0% (98/100) | 100.0% (100/100) | 1.00 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| gpt-oss-20b | autolang-loose | 99.0% (99/100) | 100.0% (100/100) | 1.00 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| gpt-oss-20b | javascript | 98.0% (98/100) | 100.0% (100/100) | 1.00 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| gpt-oss-20b | python | 100.0% (100/100) | 100.0% (100/100) | 0 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| qwen3.8-27b | autolang-strict | 99.0% (99/100) | 100.0% (100/100) | 2.00 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| qwen3.8-27b | autolang-loose | 99.0% (99/100) | 100.0% (100/100) | 2.00 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| qwen3.8-27b | javascript | 100.0% (100/100) | 100.0% (100/100) | 0 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| qwen3.8-27b | python | 100.0% (100/100) | 100.0% (100/100) | 0 | 100.0% (100/100) | 0 | 0 | 0 | 0 | 150 | 120 | 270 |
| llama-3.1-8b-instruct | autolang-strict | 74.0% (74/100) | 87.0% (87/100) | 1.77 | 87.0% (87/100) | 9 | 3 | 1 | 0 | 150 | 120 | 270 |
| llama-3.1-8b-instruct | autolang-loose | 86.0% (86/100) | 89.0% (89/100) | 3.67 | 89.0% (89/100) | 9 | 1 | 1 | 0 | 150 | 120 | 270 |
| llama-3.1-8b-instruct | javascript | 86.0% (86/100) | 86.0% (86/100) | 0 | 86.0% (86/100) | 6 | 3 | 5 | 0 | 150 | 120 | 270 |
| llama-3.1-8b-instruct | python | 88.0% (88/100) | 93.0% (93/100) | 2.80 | 93.0% (93/100) | 6 | 0 | 1 | 0 | 150 | 120 | 270 |
| qwen-2.5-7b-instruct | autolang-strict | 83.0% (83/100) | 90.0% (90/100) | 2.29 | 90.0% (90/100) | 3 | 7 | 0 | 0 | 150 | 120 | 270 |
| qwen-2.5-7b-instruct | autolang-loose | 92.0% (92/100) | 92.0% (92/100) | 0 | 92.0% (92/100) | 7 | 1 | 0 | 0 | 150 | 120 | 270 |
| qwen-2.5-7b-instruct | javascript | 86.0% (86/100) | 89.0% (89/100) | 3.00 | 89.0% (89/100) | 4 | 2 | 5 | 0 | 150 | 120 | 270 |
| qwen-2.5-7b-instruct | python | 92.0% (92/100) | 95.0% (95/100) | 1.67 | 95.0% (95/100) | 4 | 0 | 1 | 0 | 150 | 120 | 270 |
Key Architecture Findings
Predictable Error Surfaces
In JavaScript and Python environments, subtle type coercions or unhandled edge cases frequently lead to uncaught runtime exceptions during execution. Autolang transforms these unpredictable runtime failures into deterministic compilation diagnostics before execution ever begins.
High Fidelity on Smaller Models
Even on 7B–8B parameter models like Llama 3.1 8B and Qwen 2.5 7B, Autolang loose mode achieved 86%–92% Pass@1 and reached 89%–92% after retry iterations, demonstrating that capability orchestration remains dependable without requiring frontier-scale reasoning models.