AutolangDocs
Core Philosophy

The static layer absorbs the model's mistakes

In JavaScript or Python, a chain of business calls only reveals a mistake after several of them have already run. For a real workflow, those calls have already charged a card, sent an email, and written a row. Autolang checks the whole program before it runs, so a mistake stops exactly where it is, and no capability has been invoked yet.

Autolang provides language-level isolation and capability control. It replaces the general-purpose runtime that would otherwise execute inside those isolation layers.

The problem

Dynamic runtimes fail late

A dynamic language only reports the error once it reaches the wrong line. With AI-generated code, every partial run costs another model round-trip, more tokens, and a business process left half-finished.

How Autolang answers it

Compile first, execute second

Types and permissions are checked for the whole program before execution starts. A wrong argument, a wrong type, or a call outside the granted set is stopped there. No line executes.

Discovering the mistake halfway costs far more

When a host hands an agent many workflows at once, the cost of discovering an error late multiplies across them. A half-finished payment chain is worse still: the money has left, the order was never created, and nothing inside the script can undo it.

JavaScript / Python — fail late

Three calls run before anything looks wrong

  1. crm.getCustomers("enterprise") runs. One round-trip spent.
  2. payment.charge(order) runs. The money has moved.
  3. reporting.calculateMetrics(...) — misspelled. The runtime throws.

Result: two capabilities already ran and cannot be rolled back from inside the script. The host has to report the failure to the model and wait for another round-trip, or write the repair itself.

Autolang — fail early

The compiler reads the whole program first

  1. The compiler reads the entire script before anything executes.
  2. calculateMetrics is not in the granted capability set. The error returns immediately.
  3. No capability was invoked. The host has lost nothing.

Result: the error message carries the list of capabilities that do exist. The model fixes one function name on the next turn, without replaying the business process that already broke.

This is why Autolang exists. Not so AI writes better code, but so a script that is written wrong never reaches the business system at all.

The limit of this guarantee

The compiler catches names, types, and arguments. It cannot tell whether the business rule itself is right. A script that charges the wrong amount, or applies the wrong discount, compiles cleanly and runs cleanly — and it is still wrong.

Correcting that is the host's job, not the runtime's. Autolang can guarantee the blast radius of a mistake: it stays inside the granted capabilities and inside the resource budget. It cannot guarantee the intent.

Seven core principles

Every design decision in Autolang traces back to one of these.

AI is untrusted

Generated code is treated as untrusted input on every run. The runtime never assumes the model produced valid code.

Authority belongs to the host

The application owns every permission. Autolang only executes what the host registered.

Every capability is explicit

The AI never receives a database credential. It sees a named function with a declared signature and nothing else.

Execution is deterministic

The same script and the same capabilities produce the same execution path. No ambient runtime state decides the outcome.

Execution is resource bounded

Every run carries an instruction budget and a memory quota. A runaway script stops at its limit instead of consuming the host.

Default deny

Filesystem, raw network, dynamic evaluation: all denied unless the host explicitly grants them.

AI orchestrates, it does not implement

The script decides the order of business operations. The business logic itself stays in the host system.

Is Autolang right for you?

Use it when:

  • You want a wrong AI script to fail immediately, not halfway through.
  • You need to connect AI to business services without exposing credentials or connection details.
  • You plan to hand many workflows to an agent at once and want to validate all of them before anything runs.
  • You run many agents in parallel and care about per-session memory and startup time.

Skip it when:

  • You need kernel or hypervisor isolation. Autolang does not replace Docker or system virtual machines.
  • AI needs to build a large long-running application rather than execute short workflows.
  • You need to run code written by trusted humans and depend on heavy third-party libraries.

Three layers of control

Three separate boundaries: validation before execution, authority granted by the host, and resource limits at runtime.

01

The compiler absorbs mistakes

Whole-program check • Structured diagnostics
Pre-execution

Types and symbols are validated for the entire program before a single line runs. Wrong arguments, wrong types, and calls to functions that do not exist are all stopped here. When something fails, the message carries the capabilities that are actually available, so the model corrects one name on the next turn instead of guessing.

This boundary is where the absorption happens: the compiler resolves every name against the granted set, so an unknown function is a compile error rather than a call that fails halfway through. Syntax drift the model commonly produces, such as .push() or .append() for .add(), is normalised to the canonical form with a transparent warning.

02

Authority belongs to the host

Default deny • Zero ambient authority
Host security boundary

The AI never receives a credential, a socket, or an arbitrary system call. The host registers each capability explicitly; the script only orchestrates those functions. The default is deny: filesystem access, raw network, and dynamic evaluation stay blocked unless the host allows them.

03

Bounded, deterministic resources

Opcode budgets • Managed memory quota
VM boundary

Every script runs with a configured instruction budget and a memory quota. An infinite loop stops at its own limit instead of taking the host process with it.

Permission — what may AI do?The capability set the host registered. An allowlist, never an assumption.
Resource limits — how much may AI consume?The VM instruction budget and memory quota. Host-owned objects sit outside that accounting.

Capabilities in practice: the host keeps authority

The host registers business operations as typed capabilities. AI writes the orchestration and never sees a credential or a socket.

Host application (Node.js) — explicit capability registration
// Host application registers explicit capabilities compiler.registerBuiltInLibrary( "crm", ` @native("getCustomers") fun getCustomers(segment: String): Array<Customer> `, { autoImport: true }, { getCustomers: (segment) => crmService.getSegment(segment), } ); compiler.registerBuiltInLibrary( "reporting", ` @native("calculateMetrics") fun calculateMetrics(customers: Array<Customer>): Report `, { autoImport: true }, { calculateMetrics: (data) => analyticsService.calculate(data), } );
AI-generated script — orchestration only
@import("crm") @import("reporting") // AI orchestrates multiple capabilities in a single deterministic execution val enterpriseCustomers = crm.getCustomers("enterprise") val summaryReport = reporting.calculateMetrics(enterpriseCustomers) println(summaryReport)

AI → Autolang VM → Capability → Host code → Business system

If the model misspells a function name, the compiler stops the script before getCustomers ever runs. No credential is used, and no round-trip is spent asking the model what went wrong.

Strict typing does not reduce what AI can write

The reasonable objection: more type checking means more ways for a model to fail. With a minimal prompt, Autolang still reaches a first-pass success rate on par with Python and JavaScript. The reason is the language surface.

The entire prompt a code generator needs
You are a code generator for Autolang, a Kotlin-like scripting language: - Write code directly at top level (do not wrap in fun main). - Output only the code block, no explanation.

A surface models already know

Autolang uses Kotlin syntax for its foundational constructs: val, if/else, when, ?., arrayOf(). The model is not learning a new DSL.

Deliberately narrow scope

No interfaces, no sealed classes, no deep abstraction layers. Those belong to people writing libraries, not to short orchestration scripts. Fewer features means fewer ways to be wrong.

Measured: across 4 models and 350 runs per environment, Autolang recorded 0 runtime errors, with first-pass success on par with Python and JavaScript.

Full benchmark data

The gap between tool calling and schemas

Autolang sits between the two: a deterministic execution layer that allows multi-step orchestration in a single run.

Interaction layer

Tool calling, one step at a time

External queries
  • High latency: every intermediate operation needs a network round-trip.
  • Rigid flow: branching, iteration, and multi-step aggregation are awkward to express.
The missing execution gap
Orchestration layer

Autolang

Deterministic execution
  • Validated before execution: the whole script is compiled and permission-checked before the first capability call.
  • Multi-step orchestration: loops, branching, and capability composition inside one run.
  • Managed memory: VM instruction budgets and memory quotas bound every session.
Generation constraints
Constraint layer

Schemas & grammar constraints

Token validation
  • Generation constraints: forces the model to emit valid JSON or a strict grammar at the token level.
  • Complementary role: schemas govern the syntax that gets generated; Autolang provides the safe, capability-bounded place where it runs.
Vision

Where this is going

The direction is set by the same constraint: keep the surface small enough that a model can hold it, and keep the boundary strict enough that the host can rely on it.

Inference where models hesitate

Let a model declare a value without naming its type first (val list = []), with the compiler resolving the type before execution rather than trusting a guess at the call site. The same pre-execution guarantee, with fewer retries.

Capabilities that generate themselves

Derive the capability interface directly from the host registration, so the declaration a model reads is produced by the same source of truth the host enforces. No second description to keep in sync.

Where Autolang sits relative to Docker and VMs

Autolang provides language-level isolation and capability control. It does not provide OS isolation, and it does not replace Docker, KVM, or microVMs.

Autolang replaces the general-purpose runtime that would otherwise execute inside those isolation layers. Run both and you get two independent layers: infrastructure isolation at the container level, and capability control at the language level for each individual script.

Docker / microVM

Infrastructure-level isolation. Runs trusted and untrusted code alike, but every start costs time and memory.

Autolang VM

Language-level isolation for AI-generated code. Each script runs under its own instruction and memory budget and cannot reach a credential.

In short

Autolang checks the whole program before it runs. A wrong script stops at the compiler, no capability is invoked, and the business process is never left half-finished.

Authority belongs to the host. Execution is deterministic and bounded. That is the entire philosophy.