NetArena tests live network-policy repair.
Agents debug injected Kubernetes connectivity failures inside a realistic microservices environment and receive live feedback from system probes.
Deterministic infrastructure agent · Kubernetes policy
Technical report · Graph-to-policy compilation · Safety preservation · Live-probe evaluation
Benchmark — NetArena MALT Policy Benchmark
Evaluation — 2,500 dynamically generated queries
Agents debug injected Kubernetes connectivity failures inside a realistic microservices environment and receive live feedback from system probes.
A successful repair must restore the intended connection while avoiding new failures or unintended connectivity elsewhere in the service graph.
Dynamic evaluation reduces the value of memorized answers and better reflects the structure of operational network-policy work.
It is a safety-aware A2A compiler for ten public graph-operation templates and uses no model API, task IDs, seed, or stored answers.
Deterministic, safety-aware Kubernetes network-policy repair across the NetArena MALT benchmark’s public operation templates—with perfect correctness and safety at the fastest published perfect-run latency when verified.
AgentBeats ranks results by an equal-weight combination of correctness and safety, using average latency to order equally perfect results.
correctness across the published record run
safety rate while applying the required repairs
average latency, rounded from 0.05625816062092781 seconds
dynamically generated benchmark queries
NetArena evaluates the complete repair loop rather than a static policy answer.
| Dimension | What is evaluated | Why it matters | Record result |
|---|---|---|---|
| Correctness | Whether the final policy state restores the required connectivity. | A syntactically valid change is not sufficient if the intended path still fails. | 100% |
| Safety | Whether the intervention avoids creating new connectivity failures. | Operational remediation must preserve unrelated service behavior and restrictions. | 100% |
| Efficiency | Average response latency across the evaluation. | Fast policy reasoning supports interactive validation and automated control loops. | 0.056258 s |
| Generalization pressure | Dynamically generated tasks with live probe feedback. | Reduces memorization and evaluates reasoning over the presented topology. | 2,500 queries |
The registered AgentBeats profile describes a deliberately constrained, repeatable execution model.
Compiles the presented graph operation into a policy action without relying on probabilistic model output.
Treats preservation of intended restrictions and unaffected paths as a first-class objective.
Targets the benchmark’s ten public graph-operation templates rather than claiming unrestricted Kubernetes coverage.
Uses no task IDs, seed, stored answers, or model API to select its response.
The record establishes benchmark performance. Production scope must still be validated against the target cluster, network plugin, policy conventions, and operating controls.
Established by this result
Requires separate production validation
NetArena agent access
Review licensing, integration support, and deployment options for the safety-aware policy agent.
View pricing and access