Blog
Systems for building what lasts
Deep, practical guides to startups, growth, agents, inference, deploys, and the systems work behind app.nz. Every article is listenable.
Growth Engineering: how companies actually compound
Ten deep chapters spanning distribution, modern SEO and AI retrieval, positioning, activation, retention, referrals, pricing, sales, experimentation, brand, operations, and durable scale—for software startups and traditional businesses.
Learn ComfyUI in 10 runnable lessons
Install it cleanly, understand diffusion controls, build txt2img and img2img from nothing, then move through LoRAs, ControlNet, low-VRAM models, automation, and the quick-fix table.
12 runnable ComfyUI multimedia labs
Go from sound design and voice to talking characters, audio-reactive motion, coherent style systems, shader and frame transforms, structural control, and 3D asset handoffs—with upstream open-source links beside every graph.
Longevitea: building an open dataset for what survives the cup
An open research pipeline for botanical infusions: candidate chemistry, hot-water stability, evidence tiers, interaction risks, and diversity-constrained blend experiments.
Seamless textures are a topology problem: diffusion, toroidal LoRAs, verification, and a production texture bank
A paper-length engineering account of building verified tileable textures with Z-Image, exact periodic synthesis, toroidal augmentation, latent constraints, objective seam gates, cyclic embeddings, and background-priority serving.
Accelerating MiniMax H3 carefully: signed EasyCache sweeps and audiovisual gates
How app.nz tests H3 denoising caches with private signed controls, fixed-seed A/B jobs, runtime telemetry, and first/last-frame, audio, and loop quality gates.
Packaging MiniMax H3 for Cog, R2, RTX 5090, and scale to zero
A careful H3 video workflow with text, keyframes, native audio, true loops, verified R2 weights, GPU AV1, serverless pricing, and a fail-closed license gate.
Half the tokens for the same 36 tool calls: optimizing a coding agent
A model-free harness that weighs every request on the wire, the five changes that halved a 37-turn run from 1.04M to 517k tokens, and the "optimization" that made it 10% worse until the harness caught it.
A token-efficient coding agent: app agent vs Codex, measured
Head-to-head on two small repair tasks: same pass rate, 72% of the wall time, 58% of the tokens. Plus the prompt change that cut cost 2.3×, and the planner/manager loop that makes long autonomous runs safe.
Train three open piano models, then deploy one Cog anywhere
A reproducible CC0 symbolic training pipeline, listening previews, one inference image for Cog and RunPod Serverless, and one-click scale-to-zero deployment on app.nz.
Chapter 1: Growth Engineering — How the Best Startups Actually Become Big
Why breakout growth is a designed system of distribution, activation, retention, referral, and compounding—not a pile of marketing tactics.
Chapter 2: Modern SEO and AI Retrieval — Build the Source Everyone Finds
A practical content system for search intent, topical authority, programmatic pages, original evidence, internal links, and AI answer engines.
Chapter 3: Positioning and Offers — Make the Right Customer Care Quickly
Turn customer research into a sharp category, credible promise, lower-risk offer, and language that works for software and traditional businesses.
Chapter 4: Activation Engineering — Get Customers to First Value Faster
Measure the real aha moment, remove setup debt, design useful empty states, and turn the first session into a reliable path to value.
Chapter 5: Retention Systems — Build Value That Deepens Over Time
Diagnose churn by cohort, create stored value and switching benefits, design useful habits, and make customer success part of the product.
Chapter 6: Referral Loops and Networks — Let Value Travel
Engineer authentic sharing, distinguish viral effects from network effects, build partnerships, and create referral systems customers trust.
Chapter 7: Pricing and Monetization — Capture Value Without Breaking Trust
Choose a value metric, package for real segments, test willingness to pay, protect margins, and grow expansion revenue responsibly.
Chapter 8: Sales Engineering — Founder-Led Sales, Outbound, and Channels
Turn early conversations into a repeatable sales system spanning qualification, outbound, proof, pipeline, partnerships, and handoffs.
Chapter 9: The Growth Operating System — Experiments, Analytics, and Cadence
Build trustworthy measurement, prioritize experiments, avoid statistical theater, and create a weekly learning cadence that compounds.
Chapter 10: Durable Scale — Brand, Operations, Moats, and the 90-Day Plan
Combine growth loops with brand, operational capacity, capital discipline, defensibility, and a practical 90-day plan for enduring scale.
Pulsegrid: build and ship a Go WebAssembly game with app.nz CI
A complete open-source reference: deterministic Go puzzle rules, a tiny WASM bridge, portable app.yaml jobs with hybrid caches, Visualbench checks, and a static app.nz subdomain.
From ComfyUI to any Cog: open audio workflows that scale to zero
How template:name connects ComfyUI to reusable Cog deployments, when to choose open Pocket TTS versus fal audio, and how the shared serverless/pod router returns idle workers to zero.
ComfyUI lab 1: Design sound effects from first principles
Translate source, action, material, space, microphone, and timing into a reproducible SFX graph.
ComfyUI lab 2: Build loopable music beds
Separate musical role, tempo, instrumentation, density, and ending behavior instead of prompting by genre alone.
ComfyUI lab 3: Create clean narration and reusable voices
Treat speech as structured data: text, voice, pacing, pronunciation, sample rate, and provenance.
ComfyUI lab 4: Build a talking character pipeline
Generate a portrait and voice, then drive an open LiveAvatar deployment while keeping consent and identity provenance explicit.
ComfyUI lab 5: Turn sound into reactive motion
Window a waveform, measure local energy or spectrum, map it to frames, and preserve synchronization at a declared FPS.
ComfyUI lab 6: Build style as a system
Separate content, composition, light, palette, and finish; compare prompt, LoRA, IPAdapter, and deterministic transforms.
ComfyUI lab 7: Write inspectable shader looks
Use coordinate and channel math for deterministic color transforms before reaching for another diffusion pass.
ComfyUI lab 8: Prototype frame transforms safely
Understand image tensor shape, range, batch semantics, and the boundary between a prototype code block and a maintained node.
ComfyUI lab 9: Control Z-Image with line structure
Extract structure from a sketch and tune control strength separately from text guidance.
ComfyUI lab 10: Restyle without losing structure
Use a Canny signal as an explicit structural contract, then let Proteus solve appearance.
ComfyUI lab 11: Design a text-to-3D handoff
Prompt for geometry rather than beauty, generate a GLB, and inspect silhouette, topology, scale, and material assumptions.
ComfyUI lab 12: Turn a concept image into a 3D asset
Prepare an object reference that communicates shape cleanly, then validate the generated asset outside the preview renderer.
ComfyUI lesson 1: Install ComfyUI cleanly and make your first image
Pick a portable install, verify the GPU, place one checkpoint correctly, and run a known-good graph.
ComfyUI lesson 2: Diffusion, seeds, CFG, samplers, and schedulers
Learn what each setting changes by holding every other variable still.
ComfyUI lesson 3: Prompt with intent, not incantations
Layer subject, composition, light, and finish; use negatives only to correct observed failures.
ComfyUI lesson 4: Build text-to-image from a blank canvas
Wire the seven-node mental model from checkpoint to saved pixels.
ComfyUI lesson 5: Build image-to-image and master denoise
Encode a source image and use denoise to choose between a retouch and a rebuild.
ComfyUI lesson 6: Use LoRAs without losing the base model
Balance model and CLIP strength, compare with a fixed seed, and avoid adapter pileups.
ComfyUI lesson 7: Control composition with ControlNet
Preserve edges or pose while changing appearance, with a practical strength sweep.
ComfyUI lesson 8: Run modern models on modest VRAM
Choose quantization, offload, tiling, or an accelerated custom node deliberately.
ComfyUI lesson 9: Automate batches, prompt lists, REST, CLI, and MCP
Treat workflow inputs as an API contract and keep a manifest for every batch.
ComfyUI lesson 10: Debug any broken ComfyUI workflow
Start at the first red node and resolve models, nodes, VRAM, shapes, and blank outputs systematically.
Inkling: Thinking Machines’ open-weights model is built to be customized
Inside Inkling’s 975B MoE architecture, controllable reasoning, multimodal training, Apache 2.0 release, Tinker serving limits, and the new direct app.nz gateway route.
Building Polygonalize: stable low-poly video with Go, WASM, Canvas, and Three.js
How we built a free open-source image and video polygonalizer: edge-aware meshes up to 20,000 triangles, custom transparent primitives, stable video topology, Go/WASM, Canvas, Three.js, a CLI, and an app.nz serverless fallback.
Throughput on a shared GPU: fused kernels and a priority scheduler
How we serve chat, image, video, TTS, STT and a LoRA army on shared GPUs — 2× LLM tok/s (fp8+MTP), a fused multi-LoRA kernel (up to 2.4×), a 10.4× admission gate, tier-based VRAM arbitration, and cheaper building blocks (916k embeds/sec, Gemini STT).
Packaging AlayaWorld: an interactive video world model as a self-hosted Cog
Wrapping AlayaLab’s autoregressive world model (LTX-2.3-derived DiT + gemma-3-12b + Depth-Anything-3) as a Cog: camera presets, a warm engine, and an honest look at the LTX-2 Community License.
From Artificial Analysis to production: the useful eval loop
How to turn Artificial Analysis, Arena, LiveBench, SWE-bench, and open leaderboards into a runnable shortlist, workload eval, and production model decision.
Comfy models: from media evals to working GPU inference
Connect image and video benchmark evidence to hosted FLUX, SDXL, LTX, and Wan workflows, reproducible JSON, model provenance, and a global R2 weight cache.
Anatomy of an AI compute marketplace: tokens, GPU-seconds, and agent-hours on one set of rails
A deep tech dive on app.nz as a compute marketplace: 232 models across 24 providers with price-ordered self-healing failover, GPU placement that learns cold starts and arbitrages clouds under a fixed price, margins enforced as unit tests, and an agent layer that makes the whole market liquid.
The next ten years: minds as products, machine-speed markets, and why big labs are the new big tech
A ten-year thesis: AI minds with personality and skills become a product category, robots become software subscriptions, marketplaces go high-frequency and low-fee, humans and governments get augmented, and vertically integrated labs become the new big tech — with sources, and the app.nz strategy behind it.
1ms web search: gobed int8 embeddings, a self-improving index, and agent tools for flash-tier models
How app.nz search answers repeat queries in ~1ms: int8-quantized gobed embeddings at ~600 bytes/doc, a learned index fed by every search and research run, and agent tool design that cheap fast models like DeepSeek actually use well.
Five-minute training runs: capped QLoRA slices and autoresearch on RunPod
Why we cap every training run at 5 minutes: chained resumable QLoRA micro-runs, a GPU-minute budget ledger, a pod reaper, and a karpathy-style hyperparameter autoresearch loop.
Designing animation asset galleries for agents
A stable search and download contract for motion, VFX, terrain, 3D objects, REST, CLI, and MCP clients.
From reference image to indexed, downloadable 3D object
How the app.nz image-to-3D gallery connects reference images, reconstruction recipes, OBJ downloads, and production generation jobs.
Opening the app.nz motion, VFX, and terrain commons
Three pose-capture backends, retargetable character motion, configurable Three.js effects, deterministic terrain, and procedural object generators — open, searchable, and available through web, CLI, API, desktop, and MCP.
Kimi K3: the largest open model, live on the app.nz gateway
Moonshot's 2.8T-parameter Kimi K3 with a flat-priced 1M context window is available now as kimi-k3 — direct Moonshot routing, automatic failover, one metered key.
polyserve: in-process serverless that packs hundreds of apps into one process
Announcing polyserve, the open-source in-process serverless host behind app.nz: per-second billing via busy_ms and warm_s, hot-swap deploys, scale-to-zero, and autoscaling to Hetzner and RunPod GPUs.
Inside the polyserve Go host: fasthttp, unix-socket slots, and zero-drop hot swaps
How one fasthttp process serves many compiled Go apps with microsecond proxy overhead, atomic generation swaps with a 5s drain, and Postgres addons via injected DATABASE_URL.
Serving tiny ONNX models serverlessly, in-process
Why in-process beats a dedicated model server for small ONNX models: warm sessions in one asyncio host, sub-500ms cold starts, honest per-second metering, and RunPod GPU autoscaling for bigger models.
polyserve deploy docs: Go, Python, and Bun quickstarts
Per-language deploy quickstarts for polyserve spaces plus the full host control protocol: deploy, usage, health, routing, scale-to-zero, env vars, addons, and the autoscaler contract.
Interactive world models on app.nz: GPU sessions, ABot-World, and ARDY
World models need a GPU that stays up and talks over a socket. How the new session-cog protocol works, the two open-source world models you can drive today, and the harness that keeps every cog README honest with real runs.
What is an AI agent cloud?
Why app.nz is broader than a coding agent, model gateway, inference provider, or app host—and how one integrated loop changes the work of shipping AI software.
Hardening route HTML, desktop deep links, and notebook paths
A small production hardening pass: serving generated SEO HTML, keeping desktop notebook paths inside the workspace, and fixing one-pass deep-link URL encoding.
How app.nz hosted training works
The control plane behind app.nz fine-tuning: model catalogues, hardware offers, durable jobs, signed trainer specs, progress callbacks, R2 artifacts, publishing, and deploys.
Keeping training GPUs busy without wasting memory
The app.nz training worker checklist: conservative VRAM floors, bf16/TF32 defaults, efficient attention, dataloaders, direct artifact I/O, torch.compile tradeoffs, and phase-by-phase memory release.
Deploying Rust in under five seconds
Fast Rust deploys come from separating build time from release time: compile once, push a small runtime image or static bundle, then publish by swapping bytes or image refs.
How we built isolated containers for agents and CI
The app.nz isolation model: per-run Docker networks, per-step containers, service sidecars, temp Docker auth, label cleanup, timeouts, and remote machines for bigger blast-radius boundaries.
How app.nz runs GPUs on demand
Inside the GPU lifecycle: hardware catalogues with VRAM and compute capability, Cog cold starts, local GPU headroom, RunPod pods, schema introspection, metering, and idle reaping.
Why Docker cold starts are slow
A cold start is scheduling, image pull, Python import, CUDA init, weight loading, compilation, readiness, and first inference. Here is how app.nz reduces each part.
Dynamic routing across model providers
How the app.nz gateway resolves aliases, classifies auto routes, chooses BYOK or platform keys, walks fallback chains, adapts provider APIs, and meters usage from one catalogue.
Serverless vs serverful GPU inference
The GPU router behind app.nz cogs and Comfy deployments: local headroom, RunPod serverless, dedicated pods, hysteresis, single-flight cold starts, and when each tier wins.
Serving static sites from R2 and a database
The fast path behind app.nz static deploys: upload changed files, prune removed paths, index them in the database, and serve bytes directly from R2 or local storage.
Build Studio: private container builds for app.nz
How app.nz turns a Dockerfile and build context into a private registry image, with scoped Docker auth, namespaced image refs, redacted logs, and a clean release boundary.
appnz.yaml: the runtime contract
Why app.nz separates static and server runtimes with a small config file instead of guessing forever from package.json, Cargo.toml, dist folders, or Dockerfiles.
CI service containers without shared state
Inside app.nz CI: per-run Docker networks, Postgres/Redis/Mongo service aliases, per-step containers, readiness checks, timeouts, and label-based cleanup.
How coding agents become durable worker jobs
A coding-agent prompt becomes a stored task, a leased worker job, an event stream, and durable artifacts instead of a long HTTP request that disappears on failure.
Metering the model gateway without guessing
How app.nz meters token streams, image responses, media jobs, BYOK traffic, failures, and routed provider/model pairs from the same catalogue used for pricing.
BYOK provider keys in the model gateway
How app.nz lets users bring provider keys while keeping routing unified, validating providers, preferring user keys, hiding full secrets, and avoiding open-relay behavior.
Making Cog containers self-describing
How app.nz warms Cog containers, waits for real readiness, reads health and OpenAPI endpoints, and turns model schemas into typed prediction forms.
Mirroring ComfyUI models to make cold starts boring
Why app.nz mirrors popular Comfy model weights into R2 so GPU workers fetch predictable artifacts instead of relying on third-party hosts during cold starts.
RunPod serverless endpoints under the hood
How app.nz creates scale-to-zero GPU endpoints with bounded workers, queue-delay scaling, endpoint reuse, job polling, image updates, and explicit failure handling.
The lifecycle of an app.nz cloud machine
Provisioning is more than create: app.nz tracks provider ids, machine status, ready times, termination, simulated providers, and cleanup for CPU and GPU infrastructure.
Metering notebooks by the minute
How app.nz notebooks run free in the browser or on metered CPU/GPU sessions with pricing from the same endpoint, idle stop, balance stop, and durable session rows.
Queues, visibility timeouts, and dead messages
The queue primitive behind app.nz background work: receive as a lease, ack on success, retry after visibility timeout, and move poison messages to dead state.
Durable artifacts for coding agents
Why app.nz stores agent diffs, logs, previews, screenshots, and generated files as artifacts so a run remains auditable after the worker shuts down.
Cheaper GPUs by scheduling: multi-cloud arbitrage, hidden cold starts, and cogs that learn their own footprint
How app.nz cog hosting cut serving costs: community-first multi-cloud provisioning under a fixed price, serverless hedging that hides pod cold starts, and per-model learned stats (cold-start EMA, peak VRAM) that feed routing and future GPU bin-packing.
Hosted addons: Postgres with pgvector + PostGIS, and gobed GPU vector search
Attach managed services to any app or agent in one command — gobed GPU CAGRA vector search, PostgreSQL with HNSW vectors, PostGIS and graph traversal, Mongo, cache, auth, and AI monitoring — with env vars injected into every runtime.
Run cloud agent tasks on your own computers with app worker
Fire an agent task from your phone and let your desktop code it: app worker turns any machine you own into a private, user-scoped worker with local engines — our codex fork (mainline fallback), claude, cursor-agent, gemini, grok.
We open-sourced our investor room
The app.nz data room is now public pages: the 10-year vision, a full technical deep dive (architecture, ~1 ms routing, security, economics, benchmarks, roadmap, risks), all versioned in the repo.
Will it fit? An AI agent that 3D bin-packs your life into a van
A chat agent that estimates real item sizes, runs 3D bin-packing algorithms, and renders interactive three.js loading plans — with an honest verdict when it will not fit. Plus a free /api/packing/solve API.
Hosted notebooks: marimo in the browser or on cloud GPUs, billed per minute
app.nz notebooks run free on Pyodide in your browser or on per-minute CPU/GPU machines that stop themselves when idle. Plus SQLite, Parquet, and agent-trace dataset viewers with HTTP range reads.
app.nz is now an MCP server
Point Claude, Codex, or any MCP client at app.nz/mcp and your agent can run every gateway model — chat, images, embeddings — plus search the MCP directory, through one connector and one API key.
Cloud coding agents: one prompt, one pull request
Run coding agents in sandboxed cloud environments instead of on your laptop — provider-aware (Codex, Claude Code, or our runner), skill-attachable, triggerable by schedules and webhooks.
Instant static deploys: from directory to URL in one command
app.nz hosts static sites behind a URL instantly — app sites deploy uploads, diffs, and prunes a directory so a deploy is the exact state of your build, scriptable into CI or an agent run.
A searchable directory of 120+ MCP servers
What the Model Context Protocol is, how to wire a server into Claude or Codex, and a new searchable index of 120+ MCP servers filterable by category and pricing.
Test-time training is a linear attention operator in disguise
A small reproduction of a 2026 paper showing that key-value-binding TTT unrolls into a linear-attention state, plus sweeps for extra inner steps, query mismatch, gradient sign, and momentum.
An attention sink can mean two opposite things
A small reproduction of a June 2026 paper showing that the same attention stripe can be either a no-op or a broadcast channel.
The best coding models in 2026: a hands-on comparison
How to pick a coding model by the job — frontier vs agentic-mid vs fast — measured on cost per merged PR, not cost per token.
Cheap and fast: choosing budget LLMs for high-volume work
A 3-step workflow to find the cheapest model that still clears your quality bar — measured as cost per acceptable output.
How to optimize prompts with evals (LLM-as-a-judge)
Treat prompts like code: score them with an eval, then let the optimizer search for a better one against your own dataset.