RECORDED SYNTHETIC CASE / ARCHIVE INSPECTED 27 SEP 2026
Complex records.
Visible execution.
One completed live-model run on a synthetic case: 2,749,851 indexed records across 13 tables; 34 distinct seats and 114 handoff events. The source manifest digest matches the run. Analysis time excludes prior ingestion; corpus size does not mean every record entered a model prompt.
Recorded coordination: 3 lead sessions, 14 coordinator sessions, and 22 analyst sessions. Six assumption-gate events appear in the event log; they are not a count of manual approvals. These recordings do not establish the separate scenario’s 1M context or 5/0 gate configuration.
Recorded token ledger · inspect the allocation
| Role | Uncached input | Cache write | Cache read | Output |
|---|---|---|---|---|
| analyst (22) | 396 | 525,270 | 3,149,781 | 150,463 |
| coordinator (14) | 50 | 46,402 | 107,217 | 18,828 |
| lead (3) | 12 | 14,038 | 27,411 | 7,236 |
| Total | 458 | 585,710 | 3,284,409 | 176,527 |
4,047,104 recorded tokens including cached input; 84.9% of analysis/coordinator input token volume was cache reads. Cache reads are still tokens—not token elimination. The separate verifier logs contain 499,967 “tokens used” across 86 logs without an input/output/cache split; these are not folded into the provider buckets above.
Quick online-model comparison · DeepSeek pricing
Reprice the recorded analysis/coordinator token mix at DeepSeek’s published USD rates checked 27 September 2026. The original run records the alias “sonnet” and a $4.766 subtotal for these transcripts. This is a rate-card calculation, not a DeepSeek execution, a whole-run bill, or a quality comparison.
| Published model | Off-peak estimate | Peak estimate |
|---|---|---|
| DeepSeek-V4.1-Flash | $0.204 | $0.407 |
| DeepSeek-V4-Pro-0813 | $0.809 | $1.617 |
Formula: cache-read tokens × cache-hit rate + (uncached input + cache-write tokens) × cache-miss rate + output tokens × output rate, divided by one million. Verifier charges, hardware, ingestion, and service costs are excluded. Tokenizers, reasoning length, cache behavior, quality, and latency can differ in an actual run.
DeepSeek’s official API also supports tools and structured output. A fair comparison is Jianmu’s orchestration against a specified tool-enabled baseline on identical data—not “agents versus a model that cannot use tools.”
Official model specifications & rates ↗Source accounting & scope
The read-only audit checks the event sequence, final verdict corrections, dataset-manifest hash, and unique session IDs. Usage totals use only top-level session summaries; nested iterations are not added again. 38 proposals finish with 13 verified labels, 2 hypotheses, and 23 rejected labels. These are review states, not independently scored accuracy. The 34 seats are distinct roles across the run, not a simultaneous concurrency measurement.
Download the sanitized evidence summary ↓SYNTHETIC WORKLOAD / CALCULATED SCENARIO
Same 1 TB.
A different way to process it.
Project Estuary: a fictional procurement inquiry spanning payments, messages, contracts, scans, and system logs. Identify payment chains, related entities, and contradictions with source-backed findings.
A / SCENARIO BASELINE
Serial model pipeline
Repeated model reading, analysis, and review. Stages execute sequentially.
- Model token budget
- 42.00 B
- Calculated elapsed time
- 1,176.7 h
B / JIANMU SCENARIO
Hybrid execution
Scripts process records. Specialists receive selected evidence. Independent work runs in parallel.
- Model token budget
- 1.74 B
- Calculated elapsed time
- 22.5 h
Same synthetic 1 TB workload. Different execution strategies and assumed throughput. These are calculated budgets, not model test results. At the same 40,000 tokens/s rate: 301.7 h versus 22.5 h = 13.4×. Neither baseline represents DeepSeek, Claude, or all multi-agent systems.
Decimal units. Assume 120 GB of extracted text at 4 bytes/token: 30 billion token-equivalents. Binary storage bytes are never counted directly as model tokens. This page calculates the workload; it does not contain or process a 1 TB dataset.
HIERARCHICAL COORDINATION + DETERMINISTIC COMPUTE
Hundreds of tasks.
Only the necessary context.
THE VALUE / SUSTAINING MORE COMPLEX WORK
Raise the complexity.
Keep the objective intact.
The ambition is a larger span of work: more dependencies, longer investigations, and more specialist branches held together by visible state, source references, and review. Cost and time are useful measures; sustained control determines how far a project can go.
Coordination scenario assumptions, not measured requirements. One active controller uses two successive context segments—not one unlimited prompt. Automatic mode has zero scheduled pauses only within preauthorized scope; unresolved exceptions and consequential permissions can still require a person.
Compare sessions, handoffs, and human gates
| Measure | Manually coordinated model workflow | Automation-enabled MAS | Jianmu |
|---|---|---|---|
| Project-level controller | 1 active project session | 1 root role in this configuration | 1 active L1 seat |
| Controller context segments | 2 | 2 | 2 |
| Between-segment handoff | 1 manual state handoff | 1 automated handoff if configured | 1 structured continuation in this scenario |
| Task-level human touches | 1,000: dispatch + review for each package | 0 scheduled if task policies automate both | 0 scheduled in the automated task policy |
| Project human gates | 5 in this guided policy | 5 guided / 0 scheduled auto, when supported | 5 guided / 0 scheduled auto |
| Worker seats / sessions | 500 packages are work items, not 500 simultaneous seats. Actual sessions depend on reuse, retries, message sizes, context packing, and deployment limits. A 1M window alone does not determine the count. | ||
The manual baseline explicitly assumes a person dispatches and reviews each of 500 packages: 500 × 2 + 5 project gates = 1,005 workflow touches, plus one context handoff. These are configured human actions, not 1,005 mandatory security approvals or a requirement imposed by every competing product. Automated MAS can implement comparable control policies.
Context arithmetic: 500 compact results × 2,000 tokens + 400,000 tokens for instructions, decisions, and working state = 1.4M unique controller-facing tokens. With 20% headroom in a 1M window, usable capacity is 800k; ceil(1.4M / 800k) = 2 segments. This simplified budget excludes repeated transmission and assumes handoff summaries fit the reserved capacity. Worker model traffic is separate; the scenario’s 1.74B billed/processed tokens cannot be divided by 1M to infer session count.
What the five gates decide
- Authorize the objective, data scope, and boundaries.
- Approve the analysis plan and acceptance criteria.
- Review a representative sample before expanding work.
- Resolve material contradictions and exceptional decisions.
- Accept the final evidence package and authorize delivery.
Guided mode reserves these decisions for a person. The illustrative automatic policy predefines in-scope decisions and acceptance rules. It does not bypass human authority, erase verification, or guarantee no interruptions.
Long-running stability, drift control, and learning
Preserve the invariants
Carry the objective, source IDs, constraints, accepted decisions, unresolved questions, and next actions across context handoffs. Inspect omissions instead of relying on conversational memory alone.
Repair with evidence
Recompute exact operations, review interpretations independently, and recheck repaired outputs. Contradictions remain visible until resolved; an accepted label alone is not proof of truth.
Accumulate reviewed knowledge
A controlled learning loop captures corrections and reusable procedures, evaluates them on replay cases, and promotes approved versions into project memory or rules. This is governed knowledge reuse—not a claim of automatic model-weight training or guaranteed improvement.
No drift probability is assigned here: measure constraint retention, citation correctness, missed dependencies, and recovery over long tasks and repeated handoffs. More agents and a larger context window alone establish none of those rates.
One active project controller
Goal, acceptance criteria, shared state. Automatic continuation hands a compact state package to the next L1 context. A second controller can review handoff completeness.
500 logical work packages
Partition by data source, time window, and investigation question. This scenario schedules up to 32 active execution slots—not 500 simultaneous model calls.
Focused roles inside each package
Extraction, entity resolution, interpretation, and independent review activate as needed. Scripts perform exact calculations; conflicting results return for targeted repair.
Detect capacity. Set the queue.
CPU, RAM, GPU memory, storage throughput, model context, and tool availability determine scheduling width. Seat count alone does not create throughput.
Visible work. Controlled handoffs.
A task tree exposes queued, active, blocked, review, and accepted states. Shared dependencies are reserved; context handoffs retain source IDs and unresolved questions.
Seat counts and the L2/L3 arrangement above are this scenario’s configuration, not a deployment-wide concurrency promise.
AUDITABLE ARITHMETIC
Where the savings come from.
Aggregate accounting throughput across active calls, including input and output. These rates are assumptions, not hardware or model measurements.
| Stage | Serial raw-model pipeline | Context-heavy MAS | Jianmu |
|---|---|---|---|
| 01 / Ingest & normalize | 24,000 | 26,000 | 400 |
| 02 / Deep analysis | 12,000 | 16,000 | 850 |
| 03 / Tools & helper functions | 2,000 | 4,500 | 90 |
| 04 / Verify & repair | 2,500 | 6,000 | 360 |
| 05 / Handoff & delivery | 1,500 | 2,500 | 40 |
| Total | 42,000 M | 55,000 M | 1,740 M |
Inspect every stage and formula
01 / Ingest & normalize
Scripts fingerprint, deduplicate, parse, index, and aggregate source files. Models receive selected exceptions and schema samples.
Jianmu non-model critical-path allowance: 6 h02 / Deep analysis
L2 coordinators activate focused L3 specialists for each work package. Relevant windows, entity keys, and source references replace repeated whole-corpus briefings.
Jianmu non-model critical-path allowance: 1.5 h03 / Tools & helper functions
Reusable queries and scripts calculate joins, sums, time differences, and graph candidates. Model calls handle interpretation and exceptions.
Jianmu non-model critical-path allowance: 1.5 h04 / Verify & repair
Independent contexts and alternate model channels challenge candidate findings. Deterministic checks validate references and calculations; unresolved conflicts remain open.
Jianmu non-model critical-path allowance: 1 h05 / Handoff & delivery
One active L1 controller transfers bounded state to a successor context. The task tree retains dependencies, findings, blockers, and next actions.
Jianmu non-model critical-path allowance: 0.4 hToken reduction = 1 − 1,740 / 42,000 = 95.857%. Wall time = token total / effective tokens per second / 3,600 + non-model critical-path allowance. Raw pipeline allowance: 10 h; MAS: 12 h; Jianmu: 10.4 h. No extra multiplier is applied for 32 slots: their effect is already represented by aggregate throughput.
The raw baseline re-reads extracted material with additional analysis/review passes. The illustrative MAS adds repeated shared context and review messages. Jianmu’s 1.74B budget assumes selective evidence routing and script execution. These allocations are chosen scenario assumptions—not tokenizer logs or measured extraction ratios. Optimized retrieval-based or tool-using systems can narrow the difference. No model vendor or MAS framework is represented by these baselines.
Download calculation inputs ↓QUALITY IS MORE THAN AGREEMENT
Challenge the finding.
Keep its evidence.
Mechanical checks
Recalculate amounts, check time boundaries, verify source IDs, and replay joins. A missing reference blocks acceptance; an unsupported narrative remains unresolved.
Independent review
A separate context and model channel inspect the evidence and counterexamples. Model agreement alone does not establish a fact; source and rule checks remain required.
A defined quality gate
For this scenario, every accepted finding must carry inspectable source references and a recorded review decision. This is an acceptance rule, not a claim of 100% factual accuracy.
SEPARATE ARCHIVED SYNTHETIC CASE / REPORTED IN THE SUPPLIED PARTNER BRIEF
The September 27, 2026 Partner v2 brief reports 1,817 events, 34 distinct seats, and two verdict corrections. These are final archive labels, not independently scored accuracy or simultaneous concurrency. This separate archive is not the 1 TB calculation above.
For a comparative accuracy result, score accepted findings against an independent ground truth and report precision, recall, source-reference correctness, and repeated-run variability. Lower token use alone establishes none of those metrics.