建木 JIANMUCASE ANALYSIS EXPERT简体中文

RECORDED SYNTHETIC CASE / ARCHIVE INSPECTED 27 SEP 2026

Complex records.
Visible execution.

2.75Mrecords in the indexed corpus
39.8 minrecorded analysis-run elapsed time
1,255recorded query events

One completed live-model run on a synthetic case: 2,749,851 indexed records across 13 tables; 34 distinct seats and 114 handoff events. The source manifest digest matches the run. Analysis time excludes prior ingestion; corpus size does not mean every record entered a model prompt.

Recorded coordination: 3 lead sessions, 14 coordinator sessions, and 22 analyst sessions. Six assumption-gate events appear in the event log; they are not a count of manual approvals. These recordings do not establish the separate scenario’s 1M context or 5/0 gate configuration.

Recorded token ledger · inspect the allocation
39 analysis/coordinator transcripts · provider usage buckets
RoleUncached inputCache writeCache readOutput
analyst (22)396525,2703,149,781150,463
coordinator (14)5046,402107,21718,828
lead (3)1214,03827,4117,236
Total458585,7103,284,409176,527

4,047,104 recorded tokens including cached input; 84.9% of analysis/coordinator input token volume was cache reads. Cache reads are still tokens—not token elimination. The separate verifier logs contain 499,967 “tokens used” across 86 logs without an input/output/cache split; these are not folded into the provider buckets above.

Quick online-model comparison · DeepSeek pricing

Reprice the recorded analysis/coordinator token mix at DeepSeek’s published USD rates checked 27 September 2026. The original run records the alias “sonnet” and a $4.766 subtotal for these transcripts. This is a rate-card calculation, not a DeepSeek execution, a whole-run bill, or a quality comparison.

Same recorded token mix · assumed same cache hits and output count
Published modelOff-peak estimatePeak estimate
DeepSeek-V4.1-Flash$0.204$0.407
DeepSeek-V4-Pro-0813$0.809$1.617

Formula: cache-read tokens × cache-hit rate + (uncached input + cache-write tokens) × cache-miss rate + output tokens × output rate, divided by one million. Verifier charges, hardware, ingestion, and service costs are excluded. Tokenizers, reasoning length, cache behavior, quality, and latency can differ in an actual run.

DeepSeek’s official API also supports tools and structured output. A fair comparison is Jianmu’s orchestration against a specified tool-enabled baseline on identical data—not “agents versus a model that cannot use tools.”

Official model specifications & rates ↗
Source accounting & scope

The read-only audit checks the event sequence, final verdict corrections, dataset-manifest hash, and unique session IDs. Usage totals use only top-level session summaries; nested iterations are not added again. 38 proposals finish with 13 verified labels, 2 hypotheses, and 23 rejected labels. These are review states, not independently scored accuracy. The 34 seats are distinct roles across the run, not a simultaneous concurrency measurement.

Download the sanitized evidence summary ↓

SYNTHETIC WORKLOAD / CALCULATED SCENARIO

Same 1 TB.
A different way to process it.

Project Estuary: a fictional procurement inquiry spanning payments, messages, contracts, scans, and system logs. Identify payment chains, related entities, and contradictions with source-backed findings.

A / SCENARIO BASELINE

Serial model pipeline

Repeated model reading, analysis, and review. Stages execute sequentially.

Model token budget
42.00 B
Calculated elapsed time
1,176.7 h
10,000 effective tokens/s + 10 h non-model work

B / JIANMU SCENARIO

Hybrid execution

Scripts process records. Specialists receive selected evidence. Independent work runs in parallel.

Model token budget
1.74 B
Calculated elapsed time
22.5 h
40,000 effective tokens/s + 10.4 h non-model work
95.9%fewer tokens: 42.00B → 1.74B52.3×projected speed: 1,176.7 h → 22.5 h

Same synthetic 1 TB workload. Different execution strategies and assumed throughput. These are calculated budgets, not model test results. At the same 40,000 tokens/s rate: 301.7 h versus 22.5 h = 13.4×. Neither baseline represents DeepSeek, Claude, or all multi-agent systems.

340 GBTables & transactions
260 GBScans & documents
220 GBMessages & attachments
180 GBLogs & exports

Decimal units. Assume 120 GB of extracted text at 4 bytes/token: 30 billion token-equivalents. Binary storage bytes are never counted directly as model tokens. This page calculates the workload; it does not contain or process a 1 TB dataset.