Investors

Technical deep dive

Architecture, routing, security, economics, benchmarks, roadmap, and risks. Where a number is a design target rather than a measurement, it says so.

01Executive summary

app.nz is an integrated agent cloud: hosted git repos, coding agents, an OpenAI-compatible gateway routing 20+ model providers, GPU inference and training, app and site hosting, task queues and schedulers, web and paper search, and twelve editors. Everything bills to one prepaid credit balance under one API key.

Two claims. Agent traffic prefers integrated platforms, because one agent session touches repos, models, media, compute, and deploys within minutes. And a platform that sees the whole loop can optimize it, routing each request to the best model and the cheapest capable GPU in about a millisecond. That routing is where the margin comes from.

20+

model providers behind one OpenAI-compatible API

12

AI-native editors and studios sharing one asset store

$0.228/hr

GPU compute entry price, per-second billed

1

credit balance across agents, gateway, GPUs, hosting

02Product surface map

Every row is live.

SurfaceWhat it doesWhere
Repos & PRsHosted git, branch previews, review queues, CI repair/repos
Coding agentsBranch, edit, test, open PRs autonomously/agent
GatewayOpenAI-compatible routing across 20+ providers/gateway
ModelsCatalog, pricing, comparison, playground/models
Cog StudioScale-to-zero GPU endpoints from Cog containers/cogs
TrainingLoRA and full fine-tuning on per-second GPUs/training
Build Studio & CIContainer builds, registry, CI control room/builds
Hosting & sitesStatic sites, containers, workers, HTTPS subdomains/deploys
Queues & schedulersJSON job queues, cron agents, auto-agents/queues
SearchWeb search, deep research, 200M+ papers/deep-research
EditorsVideo, audio, image, vector, slides, sheets, docs, 3D, animation, whiteboard/studio
Skills & assistantsReusable agent skills, hosted assistants, characters/skills
Datasets & notebooksHosted datasets, notebooks, dashboards/datasets

Each product is a tool an agent can call, so the platform gets more useful with every tool on one account. Products we run on this stack: app.nz/papers, readingtime.app.nz, helix.app.nz, gpubrain.app.nz.

03System architecture

Control plane

A single account system owns identity, API keys, prepaid credits, spend caps, and permissions. Every product, from a gateway call to an agent minute to a GPU second, meters into the same ledger. One ledger is what makes agent spending governable and bundling possible.

Execution planes

Work runs on three planes. The request plane (gateway, search) is latency-critical, stateless, and horizontally scaled. The job plane (agents, builds, training, media renders) is queue-fed, checkpointable, and scheduled onto the cheapest capable capacity. The serving plane (deploys, hosted models, sites) is long-lived, health-checked, and migratable between providers.

Provider abstraction

Beneath all three planes sits a uniform provider layer: every model API, GPU vendor, and hosting substrate is wrapped in the same interface with live health, price, and latency telemetry. That layer is what makes the routing in chapters 04 and 05 possible, and keeps us out of any one vendor's walled garden.

Frontend

The site is a server-rendered React app with prerendered marketing routes, lazy-loaded studio surfaces, and a screenshot harness (VisualBench) that captures every page on every change. Agents use the same harness for UI review.

04Model routing & the gateway

The gateway is an OpenAI-compatible API in front of OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, Groq, Together, Fireworks, NVIDIA, OpenRouter, Fal, Netwrck, MiniMax, Z.AI, Exa, and more. Customers point an existing SDK at us, then either pin a provider model or use an auto-routed lane:

LaneOptimized forTypical use
app/autofrontier qualitygeneral chat, planning, mixed workloads
app/auto-codeagentic codingrepo edits, PR repair, migrations
app/auto-fastlatencyrealtime UX, interactive tools
app/search-deepgroundingweb, papers, citations, extraction
app/auto-image · auto-video · music · 3dmediageneration routed across media providers

Routing inputs: task lane, live provider health, published and negotiated prices, measured latency distributions, and per-customer preferences (pin, exclude, data locality). Fallback is automatic, so a provider outage degrades to the next-best provider. Fusion uses the same machinery to run a panel of models against one prompt, judge the outputs, and synthesize a stronger answer.

Why routing compounds

Any proxy can forward requests. The asset is the telemetry. Every request adds to our price, latency, and quality data across every provider, which improves the next decision, which wins more traffic. Models leapfrog each other monthly, so the router gets more valuable with each change.

05Auto-optimizing infrastructure

This is the technical core of the ten-year plan: infrastructure decisions made per request by a learned router.

The ~1 ms decision budget

A router has to be invisible: its decision time must vanish against the tens of milliseconds of network and the seconds of inference around it. Our budget is about one millisecond. That rules out calling a model to choose a model. The decision has to be a lookup: a static embedding of the request, a nearest-neighbour check against a routing index, and an arg-max over live price, latency, and health scores, all in memory on CPU. Our routing research (static-embedding kNN routers trained on model-comparison outcomes) targets that shape: quality within a point or two of an LLM-judge router at three to four orders of magnitude lower decision cost. The millisecond is a design target.

GPU placement across any provider

The same applies below the model layer. A GPU job, whether an inference burst, a training run, or a render, carries requirements (VRAM, interconnect, region, deadline) and the placer solves for the cheapest capable capacity across every substrate we can reach: serverless GPU (scale-to-zero, per-second billing, cold-start cost), hosted pods (warm, predictable, reservation economics), and spot capacity (cheapest, interruptible, needs checkpointing). Job-plane workloads checkpoint, so the placer can chase price across providers mid-workload.

SubstrateEconomicsPlaced when
Serverless GPUper-second, scale-to-zero, cold-start penaltybursty inference, low duty cycle
Hosted podsreserved-rate, always-warmsteady traffic, latency floors
Spot / preemptibledeepest discount, interruptiblecheckpointable training, batch renders
On-prem / BYOcustomer capex, zero marginaldata locality, sovereign requirements

The flywheel

Routing telemetry improves placement. Placement volume improves capacity pricing. Better pricing wins more traffic. More traffic improves telemetry. Each turn is margin from making better decisions than customers would make by hand, then splitting the savings.

06GPU compute & training

Cog Studio deploys any Cog container as a scale-to-zero HTTPS endpoint with generated input forms, prediction logs, and per-second billing from $0.228/hr. Build Studio builds images on platform workers and pushes to a private registry; the CI control room watches external pipelines (GitHub Actions and friends) and lets agents repair failures. Training runs LoRA and full fine-tunes on per-second GPUs and deploys the resulting weights to dedicated endpoints behind the same gateway key.

Kernel work matters too. Our custom Triton and CuteDSL kernels for forecasting and diffusion make owned capacity cheaper per token than rented capacity, which widens the gap the placer exploits.

07Agent runtime

A coding agent run: clone into an isolated workspace, plan against the task, edit, run tests and linters, capture VisualBench screenshots for UI changes, and open a PR with logs attached. Agents hold scoped credentials (repo-scoped tokens, per-run budgets). The whole trace, every command, diff, and model call, is replayable for review.

Around the core runner: schedulers run agents on cron; auto-agents trigger on events (failing CI, new issue, queue depth); queues feed fleets of workers with retries, dead-letter, and spend caps; and skills are versioned markdown capabilities any agent can load. There are several hundred in a searchable library, usable on our runtime or copied to any other.

08Long-running agents

The next runtime primitive is the agent that keeps running. Each requirement is a product:

RequirementWhat it meansStatus
Durable memorycontext that survives restarts and model swaps (gpubrain lineage)live
Resumable executioncheckpoint/replay so a crash or migration loses nothingbuilding
Standing budgetsmonthly spend envelopes with hard caps and alertslive
Escalation ruleswhen to act, when to page a human, when to stopbuilding
Observabilityevery action logged, diffable, attributablelive
Kill switchowner can always pause or revoke instantlylive

Target workloads: dependency and CVE patrol, p99 defense (renegotiate a service's infra nightly through the placer), content pipelines that write, illustrate, and publish continuously, and codebase gardening, the long tail of refactors no team ever schedules. Commercially, a long-running agent is a subscription that grows with the estate it manages and stays while it visibly earns its budget.

09Software that maintains itself

Self-organizing software composes everything above into one loop: telemetry, detection, proposal, proof, gated merge, deploy, back to telemetry. A system that notices its own regression, writes the fix, proves it with tests and benchmarks, and ships it through a review gate a person configured once.

What exists today

Each organ of the loop is a shipped product: repos and PRs (proposal), CI and VisualBench (proof), auto-agents (detection), deploys with previews (shipping), budgets and permissions (governance). We run early versions of the loop on our own estate, where agents file and fix real issues on the platform that hosts them.

What has to be true

Three hard problems stand between here and the category: verification strong enough to trust, goals that stay stable over long horizons, and review gates that stay meaningful when proposals arrive faster than humans read. These are platform problems, solved with better proof machinery and better gates as much as better models, which is why a platform company gets to own the answer.

10Editors for every medium

Twelve editors and studios on one substrate: video, audio, image (raster), vector, slides, sheets, docs, 3D, animation (AnimFlow), whiteboard, notebooks, and dashboards. Three design rules make them a platform rather than a bundle:

  • Agent-operable. Every editor operation is exposed as a callable tool, so an agent can cut a video or restyle a deck the same way it edits code.
  • One asset store. A generated image is immediately a video layer, a slide asset, a 3D texture source. No exports between products.
  • One balance. Every render, generation, and model call in every editor meters into the same prepaid credits as the gateway and agents.

Media generation routes like text does: images across GPT Image, FLUX, SD3, ZImage; video across Sora, Hailuo, Seedance, Wan, LTX; music, speech, and 3D across their own provider panels. Same key, same fallback, same usage logs.

11Data plane: repos & storage

Hosted git repos with branch previews, PR review queues, and GitHub-style URLs; model storage keeping weights, LoRAs, and datasets next to the workers that use them; hosted datasets and notebooks for analysis; and artifacts as the shared library every editor and agent reads and writes. Data locality is a routing input: workloads follow their data, and customers can pin regions or providers for sovereignty.

12Security

Standing practices, live today (details on /security):

  • Workload isolation. Agent runs and builds execute in isolated workspaces with scoped, short-lived credentials. Account-wide keys never enter a run.
  • Secret hygiene. Secrets are injected at runtime and scrubbed from build and CI logs, so public repos can run secure CI without leaking (edukids, our open-source flagship, runs this way).
  • Spend governance. Per-key and per-agent budgets with hard caps. Financial blast radius is a first-class control.
  • Auditability. Agent traces, gateway usage logs, and deploy histories are replayable and attributable to a key.
  • Provider containment. Customers can exclude providers or pin data locality per key; the gateway enforces it at routing time.
  • Hardened hosts. Default-deny firewalls, integrity monitoring, and alerting on the machines we operate.

We are early on formal certification. SOC 2 is on the roadmap (chapter 15). Until then, the code is public and the traces are replayable.

13Economics

How money flows

Prepaid credits are the unit of everything: customers top up a balance (or hold Pro/Ultra/Max plans that bundle monthly credits, reserved agent machines, and quotas), and every product meters against it. Prepay is both working capital and the mechanism that makes autonomous spending safe to allow.

Margin sources

SourceMechanism
Gateway spreadrouting to the cheapest capable provider; spread widens as the router learns
GPU arbitrageserverless/pod/spot placement plus owned-kernel efficiency vs rented list price
Plansbundled credits and reserved capacity with predictable utilization
Search & toolspriced per unit (e.g. papers $1/1k searches, web from $0.0077/request) over wholesale
Hostingscale-to-zero density: many idle apps per machine

Cost structure

The placer that saves customers money runs our fleet too. Scale-to-zero keeps idle surface close to free. Agents do work that would otherwise be hires. Small team, wide surface, usage-based revenue.

Under NDA

Revenue, cohort, and pipeline data are shared under NDA during a raise. Pricing and margin mechanics are public above and on /pricing.

14Benchmarks & measurement

A number is either measured and reproducible, or it is labelled a design target. Current instruments:

  • Routing evals. Static-embedding kNN routers scored against LLM-judge baselines on model-comparison datasets, giving quality versus decision-cost curves, rerun whenever the model catalog shifts. Target: judge-level lane assignment inside the ~1 ms budget (design target, tracked in the open research repo).
  • Provider telemetry. Continuous latency, error rate, and price sampling across all 20+ providers. This is the dataset the router consumes and the basis for the catalog's live pricing.
  • VisualBench. Full-page screenshot sweeps of every route on every change, desktop and mobile. UI regression as a benchmark suite (browse it at /visualbench).
  • Kernel benchmarks. Owned-kernel inference speedups (for example our forecasting-kernel work) measured against reference implementations before any capacity-cost claim is made.
  • Agent evals. Task-completion and PR-acceptance rates on our own estate. This gates what we claim agents can do.

15Roadmap

HorizonShipping
Now to 6 monthslearned lane routing in production; placer v1 across serverless/pod/spot; run-forever memory + escalation primitives; editor tool-surface completion; SOC 2 groundwork
6 to 18 monthscross-provider pod migration in GA; per-customer routing policies (cost/latency/quality dials); agent fleets over queues at scale; self-repair loop (auto-agent → fix → gated merge) as a product; certification
18 to 36 monthsrun-forever agents GA with standing budgets and audited autonomy; self-organizing estate management for customer codebases; routing index licensed as its own product
3 to 10 yearsthe vision doc, executed: the default substrate where agent-built software lives, optimizes itself, and is governed by the people who own it

When priorities move, this page changes.

16Risks

  • Platform compression. Model labs bundle more of the loop themselves. Mitigation: multi-provider neutrality is our product. Labs are unlikely to route to competitors, and customers increasingly refuse single-lab lock-in.
  • Breadth vs. depth. A wide surface risks shallow products. Mitigation: shared substrate (one ledger, one asset store, one runtime) means each surface is thin UI over deep shared infrastructure. Using it ourselves exposes shallow spots fast.
  • Provider dependence. Upstream price or ToS shifts. Mitigation: 20+ providers, live re-routing, owned capacity at the margin.
  • Autonomy incidents. An agent that does damage is the category's reputational risk. Mitigation: budgets, gates, kill switches, and replayable traces are shipped defaults, and we sell the governance as hard as the autonomy.
  • Small-team execution risk. Real. We compound with model capability and agent leverage rather than headcount.

17How to check us out

  1. Create an account and run the loop: point an agent at a repo, watch the PR arrive, check the trace.
  2. Point an OpenAI SDK at the gateway, call app/auto, and compare the usage log against any provider's list price.
  3. Read the open code: the platform's public repos, the open-source flagship (edukids), and the routing research.
  4. Browse VisualBench, the screenshot record of every surface, including its failures.
  5. Then ask for the private annex (financials, cohorts, cap table) under NDA.

It is all running, and you can have an account today.

Talk to us

Try the product, then email the founder. Financials and cap table are shared under NDA.