See the real constraint
A rover waiting for a crew slot is different from a rover stuck in traffic. Resource reservations are different from consumption. Start with an observation whose coverage is explicit.
ENGINE V1 · FIELD REPORT 03
Observe what happens. Test what might help.
Remember what actually works.
Moon is the first world. Traffic is one skill. The larger project is a civilization that accumulates competence—and an engine other worlds can use.
Enter the learning lab ↘The standalone engine, measured transport gym, typed fact verifier and M3 adapter are implemented. The research ladder and live colony integration below are the next design phase. This report does not change the running game.
01 / THE THESIS
More mind should unlock better questions, richer experiments and shared discoveries. Factories then reproduce the capabilities those discoveries make possible. The compounding resource is reliable knowledge.
A rover waiting for a crew slot is different from a rover stuck in traffic. Resource reservations are different from consumption. Start with an observation whose coverage is explicit.
More robots, a graded road or a tunnel can solve different problems. Test candidates under the same conditions, including an unchanged baseline.
Store the conditions, intervention and measured outcome. A beautiful idea that makes throughput worse is valuable evidence too. The next decision should remember it.
02 / INSIDE ONE LEARNING CYCLE
OBSERVATION / VERSIONED FACTS
Capture an immutable snapshot with units, entity references, ruleset and time. Name missing history. A model sees this bounded view, not your credentials or arbitrary world controls.
tick 138986 crew.total 16 crew.cap 10 mind.free 13.5 traffic.history unknown
Example from a captured Moon observation. A snapshot cannot prove a traffic jam.
03 / PLAY WITH THE ENGINE
Sometimes the best-looking intervention makes a queue worse. Give the colony a problem, compare the alternatives, then let the next decision consult its measured memory.
Enable JavaScript to run the shared simulator. In the default case, a road completes 8 tasks; the unchanged layout completes 4 and the exclusive tunnel completes 2.
THE NEXT DECISION
Without measured memory: keep the current layout.
Run an experiment to retain evidence on this device.
Model boundary: this is an abstract queue simulator, sharing the engine’s actual JavaScript. It does not simulate Moon terrain, collisions or current game balance. A tunnel reserves one line for the entire modeled round trip. Score = completed tasks per minute − 0.002 × intervention cost. Unfinished tasks are always shown; wait statistics cover completed tasks only. Changing world labels demonstrates adapter reuse, not proof of transfer to real warehouses.
04 / FACTS BEFORE FLUENCY
Our early model trial saw an incomplete freight summary and concluded there were no rock deliveries. The full captured state contained 15 rock packets. That failure changed the design.
The engine now checks exact typed values and entity references against the observation. Recommendations must cite their required facts. A successful check proves those claims match the supplied snapshot; it does not prove the snapshot is complete or the recommendation will work.
05 / A RESEARCH ARC FOR MIND
A proposed ladder for the game: observatories first, then comparison, experiments and cooperation. Research grants a capability; supported nodes and free mind determine whether it can operate right now.
Illustrative allocation: 3 base mind + 4 per supported node, less colony operations. All research is assumed unlocked in this demonstration. Node requirements, work costs and game research are a proposal; the standalone engine implements the gates and reservations. Its host must supply authoritative capacity. Provider token budgets are a separate limit.
Keep the colony alive and moving before reserving mind for analysis. A host capacity update interrupts a learning job that no longer fits.
Losing a node reduces the ability to run new work. It should not erase a proven design or experiment already retained.
Future research should require evidence: instrument a route, compare alternatives, reproduce a result, then share a verified capability.
06 / ONE ENGINE, MANY SKILLS
Distinguish active crew limits, mind pressure and missing route history. Recommend the next observable check.
Separate stock, reserved deliveries and unavailable consumption history. Include rock freight and actual recipe facts.
Compare candidate layouts in a deterministic gym. Retain measured wins and losses; consult them on the next decision.
Predict service demand, place spares and protect recovery capacity. Evaluate stranded time and material cost.
Explore approved parameters for throughput, power and heat. Promote tested designs into reusable factory recipes.
Publish reproducible findings with conditions and attribution. Neighbors validate a discovery before depending on it.
07 / WHAT WE ACTUALLY TESTED
This is an implementation report and a small functional trial—not a claim of general autonomous intelligence. The useful proof is concrete: a bad intervention can be measured, remembered and avoided.
| Trial | Model | Accepted | What it established |
|---|---|---|---|
| Initial prose prototype | M2.7 · historical | 4 / 6 format passes | Correct candidate selection could still accompany invented or overconfident explanations. |
| First typed protocol | M2.7 · historical | 1 / 4 | Exact claims passed, but three proposals returned a null candidate. The verifier rejected them. |
| Corrected typed protocol | M2.7 · historical | 4 / 4 | Explicit candidate enumeration resolved those four cases. Missing history produced abstention. |
| M3 initial typed trial | M3 | 2 / 4 | Two answers serialized some typed values as strings. Both were rejected. |
| M3 fact-specific schema | M3 | 2 / 4 | More explicit schema still produced two type mismatches. The independent verifier held. |
| Current provider validation | M3 | 4 / 4 | All four M3 cases accepted; all 28 claims matched their snapshots. Missing history produced abstention; measured memory selected the road. JSON-text transport is decoded once before exact typed verification. This small trial is not an accuracy guarantee. |
Tests cover restart recovery, canceled work, timeouts, quotas, stale observations, duplicate requests, ownership, capacity across SQLite connections, exact facts and loss of mind support.
The gym retains negative outcomes, reloads memory after restart and changes its next choice. An additional seed checks repeatability. Memory applies only to matching conditions; identical experiments are not counted as new evidence.
Download sanitized evidence ↗ / Read the detailed technical report ↗
08 / THE REUSABLE CORE
Versioned facts, coverage, units, candidates and authoritative support.
Persistent jobs, bounded providers, independent fact checks and audit.
Trusted simulation or future controlled trials. Outcomes scoped to their conditions.
QUEUED → ANALYZING → ANALYZED → COMPLETE
Alternative terminal states: REJECTED · FAILED · INTERRUPTED · CANCELLED
The provider proposes a candidate. It receives no game-write tool. Trusted code validates the response and a registered evaluator measures outcomes. The current Moon advisers stop after analysis; only the transport gym has an evaluator.
The host is responsible for authentication, authorization, current capacity and research unlocks. Library owner IDs are isolation keys, not proof of identity. A future live executor must add action permissions, previews, limits, monitoring and rollback criteria.
09 / RUN IT YOURSELF
The standalone starter uses Node 24 and built-in SQLite. The demo needs no API key, no package installation and no running game.
Download the engine ↓Includes source, tests, CLI and integration guide. MiniMax-M3 is the only enabled remote model.
unzip mind-engine-starter.zip cd mind-engine-starter node tests/mind-engine.test.js node scripts/mind-engine.mjs demo \ --db ./private/gym.sqlite \ --out ./private/demo.json
Typed facts, persistent jobs, budgets, verified advice, measured gym and scoped memory.
Actual history, research unlocks, shared mind allocation, player-visible evidence and bounded jobs.
Player-approved experiments, tested machine designs and reproducible federation discoveries.
10 / RESEARCH CONTEXT
Learning from retained feedback has research precedent. Reflexion explores language feedback and episodic memory without changing model weights. Voyager demonstrates a curriculum and reusable skill library in Minecraft. They inform the direction; neither validates Moon’s results or this implementation.
The current adapter uses MiniMax-M3 through the documented compatible chat interface. Tools here return an analysis to a verifier; they are not executed as world actions.
THE LONG ARC
That is how a civilization accumulates competence.
One measured result at a time.