AutolangDocs
Empirical Benchmarks

AI Code Generation Benchmarks

Empirical comparison evaluating large language models generating orchestration code across Autolang (Strict and Loose modes), JavaScript, and Python. The results evaluate first-pass success rate, error categorization, retry behavior, and runtime failure rates.

Suite v1 · 400 tasks per environment|Models 4 · 7B–27B, open-weights and commercial|Host Apple M1 Pro · npm/Wasm build|Run 2026-09-30

Figures are specific to this run. First-pass rates depend on the model as much as on the language — treat them as one measurement point, not a standing guarantee.

Runtime Reliability

0.25% Runtime Failures

Across all tested models and 400 runs per environment, Autolang recorded only 1 runtime error in each mode (0.25%), compared to 10 runtime crashes in JavaScript.

Accuracy After Retry

95.3%

In loose mode, Autolang achieved 95.3% accuracy after retry (381/400), outperforming JavaScript (93.8%) and exceeding its initial pass rate with targeted corrections.

First-Pass Accuracy

94.0% Pass@1

Autolang loose mode reached a 94.0% zero-shot first-pass success rate (376/400), higher than JavaScript (92.5%) and approaching Python (95.0%).

Overall Environment Comparison

Aggregated results across all test suites, evaluating overall reliability, error distribution, token usage, and execution latency.

EnvironmentPass@1Pass After RetryAvg RetryNominal RateLogic ErrCompile ErrRuntime ErrInfra ErrAvg TokensAvg Latency
autolang-strict88.5% (354/400)94.3% (377/400)1.8794.3% (377/400)12101027019.9 ms
autolang-loose94.0% (376/400)95.3% (381/400)4.7595.3% (381/400)1621027019.8 ms
javascript92.5% (370/400)93.8% (375/400)2.6093.8% (375/400)1051002701.1 ms
python95.0% (380/400)97.0% (388/400)2.3897.0% (388/400)1002027038.3 ms

Runtime Reliability: Unlike JavaScript which recorded 10 uncaught runtime crashes, Autolang shifts verification to deterministic compilation and capability bounds. Scripts either compile cleanly against granted host capabilities or fail deterministically at compile time, yielding only 1 runtime failure across 400 tasks.

Strict vs Loose Modes: autolang-loose relaxes strict implicit variable and bracket closing constraints, boosting Pass@1 from 88.5% to 94.0% and total success rate after retry to 95.3%.

Detailed Model and Environment Results

Granular performance across diverse open-weights and commercial models, measuring code generation fidelity and prompt efficiency.

ModelEnvironmentPass@1Pass After RetryAvg RetryNominal RateLogicCompileRuntimeInfraPrompt TokResp TokTotal Tok
gpt-oss-20bautolang-strict98.0% (98/100)100.0% (100/100)1.00100.0% (100/100)0000150120270
gpt-oss-20bautolang-loose99.0% (99/100)100.0% (100/100)1.00100.0% (100/100)0000150120270
gpt-oss-20bjavascript98.0% (98/100)100.0% (100/100)1.00100.0% (100/100)0000150120270
gpt-oss-20bpython100.0% (100/100)100.0% (100/100)0100.0% (100/100)0000150120270
qwen3.8-27bautolang-strict99.0% (99/100)100.0% (100/100)2.00100.0% (100/100)0000150120270
qwen3.8-27bautolang-loose99.0% (99/100)100.0% (100/100)2.00100.0% (100/100)0000150120270
qwen3.8-27bjavascript100.0% (100/100)100.0% (100/100)0100.0% (100/100)0000150120270
qwen3.8-27bpython100.0% (100/100)100.0% (100/100)0100.0% (100/100)0000150120270
llama-3.1-8b-instructautolang-strict74.0% (74/100)87.0% (87/100)1.7787.0% (87/100)9310150120270
llama-3.1-8b-instructautolang-loose86.0% (86/100)89.0% (89/100)3.6789.0% (89/100)9110150120270
llama-3.1-8b-instructjavascript86.0% (86/100)86.0% (86/100)086.0% (86/100)6350150120270
llama-3.1-8b-instructpython88.0% (88/100)93.0% (93/100)2.8093.0% (93/100)6010150120270
qwen-2.5-7b-instructautolang-strict83.0% (83/100)90.0% (90/100)2.2990.0% (90/100)3700150120270
qwen-2.5-7b-instructautolang-loose92.0% (92/100)92.0% (92/100)092.0% (92/100)7100150120270
qwen-2.5-7b-instructjavascript86.0% (86/100)89.0% (89/100)3.0089.0% (89/100)4250150120270
qwen-2.5-7b-instructpython92.0% (92/100)95.0% (95/100)1.6795.0% (95/100)4010150120270

Key Architecture Findings

Predictable Error Surfaces

In JavaScript and Python environments, subtle type coercions or unhandled edge cases frequently lead to uncaught runtime exceptions during execution. Autolang transforms these unpredictable runtime failures into deterministic compilation diagnostics before execution ever begins.

High Fidelity on Smaller Models

Even on 7B–8B parameter models like Llama 3.1 8B and Qwen 2.5 7B, Autolang loose mode achieved 86%–92% Pass@1 and reached 89%–92% after retry iterations, demonstrating that capability orchestration remains dependable without requiring frontier-scale reasoning models.