Watch the swarm work.
Then measure it.
AgentSaaSy is an R&D platform for enterprise AI Agent stacks. A test harness orchestrates AI Agents through real workloads while AEQ, the Agent Efficiency Quotient, scores the architecture itself: how much business value the design delivers per token it consumes. Specs do not predict adequacy. Only measuring the model-workload pair does.
The Swarm, Live
A simple workflow, end to end: a request enters, the harness routes it to a planner, fans work out to specialist AI Agents in parallel, a cross-family judge reviews the result, and the AEQ gate issues a pre-registered verdict before anything ships. Every hop is metered.
AEQ Monitor (simulated)
--BVD units / 1K tokensRun Counters
Harness Log
Architecture Layers
Four layers do the work. One measurement plane cuts across all of them. Hover or tap a layer to inspect it.
AgentSaaSy Application Layer
Workflow definitions, business value rubrics, and the SaaS-substitution surface where AI Agent stacks replace seat-licensed software.
Harness & Orchestration Layer
The conductor. Routes requests, fans out parallel work, enforces pre-registered gate thresholds, and coordinates cross-family judging.
Agent Layer
The swarm: planner, retriever, workers, and an independent judge from a different model family than the workers it reviews.
Model Layer
Frontier APIs and quantized local models, treated as interchangeable capacity. The pair (model + workload) is what gets measured, never the spec sheet.
Layer Detail
The Stack
What actually runs underneath the demo above.
From Simple to Swarm-Scale
This page shows one workflow. The architecture is built to grow.
One workflow, fully metered
Single request pipeline with parallel fan-out, cross-family judging, and a pre-registered AEQ gate on every run.
Multi-workflow swarms
Concurrent workflows sharing the agent pool, tiered model routing, and per-workflow AEQ baselines to catch architecture waste.
AEQ Grid certification
The full 3x3x3 grid: query classes by model tiers by repeated runs, producing GREEN / YELLOW / RED verdicts for model-workload pairs.