Investors
Technical deep dive
Architecture, routing, security, economics, benchmarks, roadmap, and risks. Where a number is a design target rather than a measurement, it says so.
01Executive summary
app.nz is an integrated agent cloud: hosted git repos, coding agents, an OpenAI-compatible gateway routing 20+ model providers, GPU inference and training, app and site hosting, task queues and schedulers, web and paper search, and twelve editors. Everything bills to one prepaid credit balance under one API key.
Two claims. Agent traffic prefers integrated platforms, because one agent session touches repos, models, media, compute, and deploys within minutes. And a platform that sees the whole loop can optimize it, routing each request to the best model and the cheapest capable GPU in about a millisecond. That routing is where the margin comes from.
20+
model providers behind one OpenAI-compatible API
12
AI-native editors and studios sharing one asset store
$0.228/hr
GPU compute entry price, per-second billed
1
credit balance across agents, gateway, GPUs, hosting
02Product surface map
Every row is live.
| Surface | What it does | Where |
|---|---|---|
| Repos & PRs | Hosted git, branch previews, review queues, CI repair | /repos |
| Coding agents | Branch, edit, test, open PRs autonomously | /agent |
| Gateway | OpenAI-compatible routing across 20+ providers | /gateway |
| Models | Catalog, pricing, comparison, playground | /models |
| Cog Studio | Scale-to-zero GPU endpoints from Cog containers | /cogs |
| Training | LoRA and full fine-tuning on per-second GPUs | /training |
| Build Studio & CI | Container builds, registry, CI control room | /builds |
| Hosting & sites | Static sites, containers, workers, HTTPS subdomains | /deploys |
| Queues & schedulers | JSON job queues, cron agents, auto-agents | /queues |
| Search | Web search, deep research, 200M+ papers | /deep-research |
| Editors | Video, audio, image, vector, slides, sheets, docs, 3D, animation, whiteboard | /studio |
| Skills & assistants | Reusable agent skills, hosted assistants, characters | /skills |
| Datasets & notebooks | Hosted datasets, notebooks, dashboards | /datasets |
Each product is a tool an agent can call, so the platform gets more useful with every tool on one account. Products we run on this stack: app.nz/papers, readingtime.app.nz, helix.app.nz, gpubrain.app.nz.
03System architecture
Control plane
A single account system owns identity, API keys, prepaid credits, spend caps, and permissions. Every product, from a gateway call to an agent minute to a GPU second, meters into the same ledger. One ledger is what makes agent spending governable and bundling possible.
Execution planes
Work runs on three planes. The request plane (gateway, search) is latency-critical, stateless, and horizontally scaled. The job plane (agents, builds, training, media renders) is queue-fed, checkpointable, and scheduled onto the cheapest capable capacity. The serving plane (deploys, hosted models, sites) is long-lived, health-checked, and migratable between providers.
Provider abstraction
Beneath all three planes sits a uniform provider layer: every model API, GPU vendor, and hosting substrate is wrapped in the same interface with live health, price, and latency telemetry. That layer is what makes the routing in chapters 04 and 05 possible, and keeps us out of any one vendor's walled garden.
Frontend
The site is a server-rendered React app with prerendered marketing routes, lazy-loaded studio surfaces, and a screenshot harness (VisualBench) that captures every page on every change. Agents use the same harness for UI review.
04Model routing & the gateway
The gateway is an OpenAI-compatible API in front of OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, Groq, Together, Fireworks, NVIDIA, OpenRouter, Fal, Netwrck, MiniMax, Z.AI, Exa, and more. Customers point an existing SDK at us, then either pin a provider model or use an auto-routed lane:
| Lane | Optimized for | Typical use |
|---|---|---|
| app/auto | frontier quality | general chat, planning, mixed workloads |
| app/auto-code | agentic coding | repo edits, PR repair, migrations |
| app/auto-fast | latency | realtime UX, interactive tools |
| app/search-deep | grounding | web, papers, citations, extraction |
| app/auto-image · auto-video · music · 3d | media | generation routed across media providers |
Routing inputs: task lane, live provider health, published and negotiated prices, measured latency distributions, and per-customer preferences (pin, exclude, data locality). Fallback is automatic, so a provider outage degrades to the next-best provider. Fusion uses the same machinery to run a panel of models against one prompt, judge the outputs, and synthesize a stronger answer.
Why routing compounds
Any proxy can forward requests. The asset is the telemetry. Every request adds to our price, latency, and quality data across every provider, which improves the next decision, which wins more traffic. Models leapfrog each other monthly, so the router gets more valuable with each change.
05Auto-optimizing infrastructure
This is the technical core of the ten-year plan: infrastructure decisions made per request by a learned router.
The ~1 ms decision budget
A router has to be invisible: its decision time must vanish against the tens of milliseconds of network and the seconds of inference around it. Our budget is about one millisecond. That rules out calling a model to choose a model. The decision has to be a lookup: a static embedding of the request, a nearest-neighbour check against a routing index, and an arg-max over live price, latency, and health scores, all in memory on CPU. Our routing research (static-embedding kNN routers trained on model-comparison outcomes) targets that shape: quality within a point or two of an LLM-judge router at three to four orders of magnitude lower decision cost. The millisecond is a design target.
GPU placement across any provider
The same applies below the model layer. A GPU job, whether an inference burst, a training run, or a render, carries requirements (VRAM, interconnect, region, deadline) and the placer solves for the cheapest capable capacity across every substrate we can reach: serverless GPU (scale-to-zero, per-second billing, cold-start cost), hosted pods (warm, predictable, reservation economics), and spot capacity (cheapest, interruptible, needs checkpointing). Job-plane workloads checkpoint, so the placer can chase price across providers mid-workload.
| Substrate | Economics | Placed when |
|---|---|---|
| Serverless GPU | per-second, scale-to-zero, cold-start penalty | bursty inference, low duty cycle |
| Hosted pods | reserved-rate, always-warm | steady traffic, latency floors |
| Spot / preemptible | deepest discount, interruptible | checkpointable training, batch renders |
| On-prem / BYO | customer capex, zero marginal | data locality, sovereign requirements |
The flywheel
Routing telemetry improves placement. Placement volume improves capacity pricing. Better pricing wins more traffic. More traffic improves telemetry. Each turn is margin from making better decisions than customers would make by hand, then splitting the savings.
06GPU compute & training
Cog Studio deploys any Cog container as a scale-to-zero HTTPS endpoint with generated input forms, prediction logs, and per-second billing from $0.228/hr. Build Studio builds images on platform workers and pushes to a private registry; the CI control room watches external pipelines (GitHub Actions and friends) and lets agents repair failures. Training runs LoRA and full fine-tunes on per-second GPUs and deploys the resulting weights to dedicated endpoints behind the same gateway key.
Kernel work matters too. Our custom Triton and CuteDSL kernels for forecasting and diffusion make owned capacity cheaper per token than rented capacity, which widens the gap the placer exploits.
07Agent runtime
A coding agent run: clone into an isolated workspace, plan against the task, edit, run tests and linters, capture VisualBench screenshots for UI changes, and open a PR with logs attached. Agents hold scoped credentials (repo-scoped tokens, per-run budgets). The whole trace, every command, diff, and model call, is replayable for review.
Around the core runner: schedulers run agents on cron; auto-agents trigger on events (failing CI, new issue, queue depth); queues feed fleets of workers with retries, dead-letter, and spend caps; and skills are versioned markdown capabilities any agent can load. There are several hundred in a searchable library, usable on our runtime or copied to any other.
08Long-running agents
The next runtime primitive is the agent that keeps running. Each requirement is a product:
| Requirement | What it means | Status |
|---|---|---|
| Durable memory | context that survives restarts and model swaps (gpubrain lineage) | live |
| Resumable execution | checkpoint/replay so a crash or migration loses nothing | building |
| Standing budgets | monthly spend envelopes with hard caps and alerts | live |
| Escalation rules | when to act, when to page a human, when to stop | building |
| Observability | every action logged, diffable, attributable | live |
| Kill switch | owner can always pause or revoke instantly | live |
Target workloads: dependency and CVE patrol, p99 defense (renegotiate a service's infra nightly through the placer), content pipelines that write, illustrate, and publish continuously, and codebase gardening, the long tail of refactors no team ever schedules. Commercially, a long-running agent is a subscription that grows with the estate it manages and stays while it visibly earns its budget.
09Software that maintains itself
Self-organizing software composes everything above into one loop: telemetry, detection, proposal, proof, gated merge, deploy, back to telemetry. A system that notices its own regression, writes the fix, proves it with tests and benchmarks, and ships it through a review gate a person configured once.
What exists today
Each organ of the loop is a shipped product: repos and PRs (proposal), CI and VisualBench (proof), auto-agents (detection), deploys with previews (shipping), budgets and permissions (governance). We run early versions of the loop on our own estate, where agents file and fix real issues on the platform that hosts them.
What has to be true
Three hard problems stand between here and the category: verification strong enough to trust, goals that stay stable over long horizons, and review gates that stay meaningful when proposals arrive faster than humans read. These are platform problems, solved with better proof machinery and better gates as much as better models, which is why a platform company gets to own the answer.
10Editors for every medium
Twelve editors and studios on one substrate: video, audio, image (raster), vector, slides, sheets, docs, 3D, animation (AnimFlow), whiteboard, notebooks, and dashboards. Three design rules make them a platform rather than a bundle:
- Agent-operable. Every editor operation is exposed as a callable tool, so an agent can cut a video or restyle a deck the same way it edits code.
- One asset store. A generated image is immediately a video layer, a slide asset, a 3D texture source. No exports between products.
- One balance. Every render, generation, and model call in every editor meters into the same prepaid credits as the gateway and agents.
Media generation routes like text does: images across GPT Image, FLUX, SD3, ZImage; video across Sora, Hailuo, Seedance, Wan, LTX; music, speech, and 3D across their own provider panels. Same key, same fallback, same usage logs.
11Data plane: repos & storage
Hosted git repos with branch previews, PR review queues, and GitHub-style URLs; model storage keeping weights, LoRAs, and datasets next to the workers that use them; hosted datasets and notebooks for analysis; and artifacts as the shared library every editor and agent reads and writes. Data locality is a routing input: workloads follow their data, and customers can pin regions or providers for sovereignty.
12Security
Standing practices, live today (details on /security):
- Workload isolation. Agent runs and builds execute in isolated workspaces with scoped, short-lived credentials. Account-wide keys never enter a run.
- Secret hygiene. Secrets are injected at runtime and scrubbed from build and CI logs, so public repos can run secure CI without leaking (edukids, our open-source flagship, runs this way).
- Spend governance. Per-key and per-agent budgets with hard caps. Financial blast radius is a first-class control.
- Auditability. Agent traces, gateway usage logs, and deploy histories are replayable and attributable to a key.
- Provider containment. Customers can exclude providers or pin data locality per key; the gateway enforces it at routing time.
- Hardened hosts. Default-deny firewalls, integrity monitoring, and alerting on the machines we operate.
We are early on formal certification. SOC 2 is on the roadmap (chapter 15). Until then, the code is public and the traces are replayable.
13Economics
How money flows
Prepaid credits are the unit of everything: customers top up a balance (or hold Pro/Ultra/Max plans that bundle monthly credits, reserved agent machines, and quotas), and every product meters against it. Prepay is both working capital and the mechanism that makes autonomous spending safe to allow.
Margin sources
| Source | Mechanism |
|---|---|
| Gateway spread | routing to the cheapest capable provider; spread widens as the router learns |
| GPU arbitrage | serverless/pod/spot placement plus owned-kernel efficiency vs rented list price |
| Plans | bundled credits and reserved capacity with predictable utilization |
| Search & tools | priced per unit (e.g. papers $1/1k searches, web from $0.0077/request) over wholesale |
| Hosting | scale-to-zero density: many idle apps per machine |
Cost structure
The placer that saves customers money runs our fleet too. Scale-to-zero keeps idle surface close to free. Agents do work that would otherwise be hires. Small team, wide surface, usage-based revenue.
Under NDA
Revenue, cohort, and pipeline data are shared under NDA during a raise. Pricing and margin mechanics are public above and on /pricing.
14Benchmarks & measurement
A number is either measured and reproducible, or it is labelled a design target. Current instruments:
- Routing evals. Static-embedding kNN routers scored against LLM-judge baselines on model-comparison datasets, giving quality versus decision-cost curves, rerun whenever the model catalog shifts. Target: judge-level lane assignment inside the ~1 ms budget (design target, tracked in the open research repo).
- Provider telemetry. Continuous latency, error rate, and price sampling across all 20+ providers. This is the dataset the router consumes and the basis for the catalog's live pricing.
- VisualBench. Full-page screenshot sweeps of every route on every change, desktop and mobile. UI regression as a benchmark suite (browse it at /visualbench).
- Kernel benchmarks. Owned-kernel inference speedups (for example our forecasting-kernel work) measured against reference implementations before any capacity-cost claim is made.
- Agent evals. Task-completion and PR-acceptance rates on our own estate. This gates what we claim agents can do.
15Roadmap
| Horizon | Shipping |
|---|---|
| Now to 6 months | learned lane routing in production; placer v1 across serverless/pod/spot; run-forever memory + escalation primitives; editor tool-surface completion; SOC 2 groundwork |
| 6 to 18 months | cross-provider pod migration in GA; per-customer routing policies (cost/latency/quality dials); agent fleets over queues at scale; self-repair loop (auto-agent → fix → gated merge) as a product; certification |
| 18 to 36 months | run-forever agents GA with standing budgets and audited autonomy; self-organizing estate management for customer codebases; routing index licensed as its own product |
| 3 to 10 years | the vision doc, executed: the default substrate where agent-built software lives, optimizes itself, and is governed by the people who own it |
When priorities move, this page changes.
16Risks
- Platform compression. Model labs bundle more of the loop themselves. Mitigation: multi-provider neutrality is our product. Labs are unlikely to route to competitors, and customers increasingly refuse single-lab lock-in.
- Breadth vs. depth. A wide surface risks shallow products. Mitigation: shared substrate (one ledger, one asset store, one runtime) means each surface is thin UI over deep shared infrastructure. Using it ourselves exposes shallow spots fast.
- Provider dependence. Upstream price or ToS shifts. Mitigation: 20+ providers, live re-routing, owned capacity at the margin.
- Autonomy incidents. An agent that does damage is the category's reputational risk. Mitigation: budgets, gates, kill switches, and replayable traces are shipped defaults, and we sell the governance as hard as the autonomy.
- Small-team execution risk. Real. We compound with model capability and agent leverage rather than headcount.
17How to check us out
- Create an account and run the loop: point an agent at a repo, watch the PR arrive, check the trace.
- Point an OpenAI SDK at the gateway, call
app/auto, and compare the usage log against any provider's list price. - Read the open code: the platform's public repos, the open-source flagship (edukids), and the routing research.
- Browse VisualBench, the screenshot record of every surface, including its failures.
- Then ask for the private annex (financials, cohorts, cap table) under NDA.
It is all running, and you can have an account today.
Talk to us
Try the product, then email the founder. Financials and cap table are shared under NDA.