app.nzapp
AppsProjectsReposPullsChatIntegrationsGatewayModelsEvalsToolsDatasetsMCPDeploysPricingBlogDocsAssistantsCharactersArtMusic
Sign inStart building
Agent stack
Cloud coding agentAgents SDKIntegrationsBrowser agentMonitors & auto-agentsSchedulersAgent skillsMCP serversDeep research
Models & API
AI GatewayModel catalogModel evalsModel spacesPlaygroundText to imageImage to 3DText to 3DMusic & SFXAudio editorMedia optimizerAI art & libraryChatAPI referenceSchemaBecome a provider
Compute & hosting
DeploysAddonsPostgres hostinggobed vector searchSite hostingAnalyticsCog GPU hostingRL trainingBuilds & CIWorkersTask queuesDomainsGit hosting
Tools
AI toolsDrawDiffusion canvasLive DrawWriteSheetsArtifactsVideo studioNotebooksDatasets
Learn
DocsBlogEval guidesPrompt libraryCLIAlternativesPapersAI charactersArt gallerySecurityConsulting
Company
PricingEnterpriseSettingsBillingStatusInvestorsCreate accountTerms of ServicePrivacy Policy
app.nzapp.nz

AI agent cloud for coding, deploys, model routing, and research. Built for teams shipping software.

Built in New Zealand by App AI NZ.

Social
X / TwitterGitHubYouTube
The app.nz network
GpuBrainPapersReading TimemojojojoNetwrckText-Generator.ioCodex InfinityOpenPathsCuteDSLAI Art GeneratorAIArt-Generator.artSiteSimSimplexGenDictatorFlowWebFiddleRing.nzChatGibidyBitBankExperimentFlowEvangelerHires.nzHow.nzV5 GamesAddicting Word GamesBig Multiplayer ChessWord SmashingreWord GameMultiplication Master
© 2026 App AI NZ Ltd. All rights reserved.All systems normalTermsPrivacy

Blog

Systems for building what lasts

Deep, practical guides to startups, growth, agents, inference, deploys, and the systems work behind app.nz. Every article is listenable.

New · complete audio series

Growth Engineering: how companies actually compound

Ten deep chapters spanning distribution, modern SEO and AI retrieval, positioning, activation, retention, referrals, pricing, sales, experimentation, brand, operations, and durable scale—for software startups and traditional businesses.

Start chapter 1 Locally voiced with OmniServe10 chapters · 126 min read
New practical series

Learn ComfyUI in 10 runnable lessons

Install it cleanly, understand diffusion controls, build txt2img and img2img from nothing, then move through LoRAs, ControlNet, low-VRAM models, automation, and the quick-fix table.

Start lesson 1 Browse runnable workflows10 parts · 315 min hands-on
Creative systems collection

12 runnable ComfyUI multimedia labs

Go from sound design and voice to talking characters, audio-reactive motion, coherent style systems, shader and frame transforms, structural control, and 3D asset handoffs—with upstream open-source links beside every graph.

Start lab 1 Run all lab workflows12 labs · 440 min hands-on
Aug 10, 2026·8 min read· Listen

Longevitea: building an open dataset for what survives the cup

An open research pipeline for botanical infusions: candidate chemistry, hot-water stability, evidence tiers, interaction risks, and diversity-constrained blend experiments.

researchnutritionteaopen-sourcedatasetsPlotly
Aug 4, 2026·18 min read· Listen

Seamless textures are a topology problem: diffusion, toroidal LoRAs, verification, and a production texture bank

A paper-length engineering account of building verified tileable textures with Z-Image, exact periodic synthesis, toroidal augmentation, latent constraints, objective seam gates, cyclic embeddings, and background-priority serving.

diffusiontexturesz-imagelorainferenceresearchgpu
Aug 4, 2026·7 min read· Listen

Accelerating MiniMax H3 carefully: signed EasyCache sweeps and audiovisual gates

How app.nz tests H3 denoising caches with private signed controls, fixed-seed A/B jobs, runtime telemetry, and first/last-frame, audio, and loop quality gates.

modelsvideoperformancegpucomfyuievalsserverless
Aug 4, 2026·8 min read· Listen

Packaging MiniMax H3 for Cog, R2, RTX 5090, and scale to zero

A careful H3 video workflow with text, keyframes, native audio, true loops, verified R2 weights, GPU AV1, serverless pricing, and a fail-closed license gate.

modelsvideocogrunpodserverlessgpuopen-source
Jul 30, 2026·9 min read· Listen

Half the tokens for the same 36 tool calls: optimizing a coding agent

A model-free harness that weighs every request on the wire, the five changes that halved a 37-turn run from 1.04M to 517k tokens, and the "optimization" that made it 10% worse until the harness caught it.

agentsclitokensbenchmarksopen-sourceperformance
Jul 25, 2026·7 min read· Listen

A token-efficient coding agent: app agent vs Codex, measured

Head-to-head on two small repair tasks: same pass rate, 72% of the wall time, 58% of the tokens. Plus the prompt change that cut cost 2.3×, and the planner/manager loop that makes long autonomous runs safe.

agentsbenchmarkscliopen-sourcetokenscodex
Jul 24, 2026·9 min read· Listen

Train three open piano models, then deploy one Cog anywhere

A reproducible CC0 symbolic training pipeline, listening previews, one inference image for Cog and RunPod Serverless, and one-click scale-to-zero deployment on app.nz.

open-sourcemodelsmusicpianocogrunpodserverlesstutorial
Jul 24, 2026·15 min read· Listen

Chapter 1: Growth Engineering — How the Best Startups Actually Become Big

Why breakout growth is a designed system of distribution, activation, retention, referral, and compounding—not a pile of marketing tactics.

startupsgrowth engineeringbusinessgrowth engineering
Jul 24, 2026·13 min read· Listen

Chapter 2: Modern SEO and AI Retrieval — Build the Source Everyone Finds

A practical content system for search intent, topical authority, programmatic pages, original evidence, internal links, and AI answer engines.

startupsgrowth engineeringbusinessseo and ai retrieval
Jul 24, 2026·12 min read· Listen

Chapter 3: Positioning and Offers — Make the Right Customer Care Quickly

Turn customer research into a sharp category, credible promise, lower-risk offer, and language that works for software and traditional businesses.

startupsgrowth engineeringbusinesspositioning and offers
Jul 24, 2026·11 min read· Listen

Chapter 4: Activation Engineering — Get Customers to First Value Faster

Measure the real aha moment, remove setup debt, design useful empty states, and turn the first session into a reliable path to value.

startupsgrowth engineeringbusinessactivation and onboarding
Jul 24, 2026·12 min read· Listen

Chapter 5: Retention Systems — Build Value That Deepens Over Time

Diagnose churn by cohort, create stored value and switching benefits, design useful habits, and make customer success part of the product.

startupsgrowth engineeringbusinessretention systems
Jul 24, 2026·11 min read· Listen

Chapter 6: Referral Loops and Networks — Let Value Travel

Engineer authentic sharing, distinguish viral effects from network effects, build partnerships, and create referral systems customers trust.

startupsgrowth engineeringbusinessreferrals and networks
Jul 24, 2026·12 min read· Listen

Chapter 7: Pricing and Monetization — Capture Value Without Breaking Trust

Choose a value metric, package for real segments, test willingness to pay, protect margins, and grow expansion revenue responsibly.

startupsgrowth engineeringbusinesspricing and monetization
Jul 24, 2026·13 min read· Listen

Chapter 8: Sales Engineering — Founder-Led Sales, Outbound, and Channels

Turn early conversations into a repeatable sales system spanning qualification, outbound, proof, pipeline, partnerships, and handoffs.

startupsgrowth engineeringbusinesssales engineering
Jul 24, 2026·13 min read· Listen

Chapter 9: The Growth Operating System — Experiments, Analytics, and Cadence

Build trustworthy measurement, prioritize experiments, avoid statistical theater, and create a weekly learning cadence that compounds.

startupsgrowth engineeringbusinessgrowth operating system
Jul 24, 2026·14 min read· Listen

Chapter 10: Durable Scale — Brand, Operations, Moats, and the 90-Day Plan

Combine growth loops with brand, operational capacity, capital discipline, defensibility, and a practical 90-day plan for enduring scale.

startupsgrowth engineeringbusinessdurable scale
Jul 23, 2026·8 min read· Listen

Pulsegrid: build and ship a Go WebAssembly game with app.nz CI

A complete open-source reference: deterministic Go puzzle rules, a tiny WASM bridge, portable app.yaml jobs with hybrid caches, Visualbench checks, and a static app.nz subdomain.

golangwasmciopen-sourcegamestutorial
Jul 22, 2026·10 min read· Listen

From ComfyUI to any Cog: open audio workflows that scale to zero

How template:name connects ComfyUI to reusable Cog deployments, when to choose open Pocket TTS versus fal audio, and how the shared serverless/pod router returns idle workers to zero.

comfyuicogaudioopen-sourceserverlessgpututorial
Jul 22, 2026·8 min read· Listen

ComfyUI lab 1: Design sound effects from first principles

Translate source, action, material, space, microphone, and timing into a reproducible SFX graph.

comfycomfyuiworkflow labsoundenvelopesprompt systems
Jul 22, 2026·8 min read· Listen

ComfyUI lab 2: Build loopable music beds

Separate musical role, tempo, instrumentation, density, and ending behavior instead of prompting by genre alone.

comfycomfyuiworkflow labmusictempoloops
Jul 22, 2026·8 min read· Listen

ComfyUI lab 3: Create clean narration and reusable voices

Treat speech as structured data: text, voice, pacing, pronunciation, sample rate, and provenance.

comfycomfyuiworkflow labttsvoiceaudio
Jul 22, 2026·10 min read· Listen

ComfyUI lab 4: Build a talking character pipeline

Generate a portrait and voice, then drive an open LiveAvatar deployment while keeping consent and identity provenance explicit.

comfycomfyuiworkflow labcharacterlip syncliveavatar
Jul 22, 2026·9 min read· Listen

ComfyUI lab 5: Turn sound into reactive motion

Window a waveform, measure local energy or spectrum, map it to frames, and preserve synchronization at a declared FPS.

comfycomfyuiworkflow labreactivefftanimation
Jul 22, 2026·9 min read· Listen

ComfyUI lab 6: Build style as a system

Separate content, composition, light, palette, and finish; compare prompt, LoRA, IPAdapter, and deterministic transforms.

comfycomfyuiworkflow labstyleconditioningipadapter
Jul 22, 2026·8 min read· Listen

ComfyUI lab 7: Write inspectable shader looks

Use coordinate and channel math for deterministic color transforms before reaching for another diffusion pass.

comfycomfyuiworkflow labshaderpixelsdeterminism
Jul 22, 2026·8 min read· Listen

ComfyUI lab 8: Prototype frame transforms safely

Understand image tensor shape, range, batch semantics, and the boundary between a prototype code block and a maintained node.

comfycomfyuiworkflow labpythontensorsframes
Jul 22, 2026·9 min read· Listen

ComfyUI lab 9: Control Z-Image with line structure

Extract structure from a sketch and tune control strength separately from text guidance.

comfycomfyuiworkflow labz-imagecontrolnetcutedsl
Jul 22, 2026·8 min read· Listen

ComfyUI lab 10: Restyle without losing structure

Use a Canny signal as an explicit structural contract, then let Proteus solve appearance.

comfycomfyuiworkflow labproteuscannystyle transfer
Jul 22, 2026·8 min read· Listen

ComfyUI lab 11: Design a text-to-3D handoff

Prompt for geometry rather than beauty, generate a GLB, and inspect silhouette, topology, scale, and material assumptions.

comfycomfyuiworkflow lab3dgeometryglb
Jul 22, 2026·8 min read· Listen

ComfyUI lab 12: Turn a concept image into a 3D asset

Prepare an object reference that communicates shape cleanly, then validate the generated asset outside the preview renderer.

comfycomfyuiworkflow lab3dreferenceasset qa
Jul 22, 2026·7 min read· Listen

ComfyUI lesson 1: Install ComfyUI cleanly and make your first image

Pick a portable install, verify the GPU, place one checkpoint correctly, and run a known-good graph.

comfycomfyuitutorialinstallmodelsfirst run
Jul 22, 2026·6 min read· Listen

ComfyUI lesson 2: Diffusion, seeds, CFG, samplers, and schedulers

Learn what each setting changes by holding every other variable still.

comfycomfyuitutorialdiffusionseedcfgsampler
Jul 22, 2026·6 min read· Listen

ComfyUI lesson 3: Prompt with intent, not incantations

Layer subject, composition, light, and finish; use negatives only to correct observed failures.

comfycomfyuitutorialpromptingnegative promptexperiments
Jul 22, 2026·7 min read· Listen

ComfyUI lesson 4: Build text-to-image from a blank canvas

Wire the seven-node mental model from checkpoint to saved pixels.

comfycomfyuitutorialtext-to-imagenodesvae
Jul 22, 2026·6 min read· Listen

ComfyUI lesson 5: Build image-to-image and master denoise

Encode a source image and use denoise to choose between a retouch and a rebuild.

comfycomfyuitutorialimage-to-imagevaedenoise
Jul 22, 2026·7 min read· Listen

ComfyUI lesson 6: Use LoRAs without losing the base model

Balance model and CLIP strength, compare with a fixed seed, and avoid adapter pileups.

comfycomfyuitutoriallorastylestrength
Jul 22, 2026·6 min read· Listen

ComfyUI lesson 7: Control composition with ControlNet

Preserve edges or pose while changing appearance, with a practical strength sweep.

comfycomfyuitutorialcontrolnetcannycomposition
Jul 22, 2026·8 min read· Listen

ComfyUI lesson 8: Run modern models on modest VRAM

Choose quantization, offload, tiling, or an accelerated custom node deliberately.

comfycomfyuitutorialfp8ggufvramcutedsl
Jul 22, 2026·8 min read· Listen

ComfyUI lesson 9: Automate batches, prompt lists, REST, CLI, and MCP

Treat workflow inputs as an API contract and keep a manifest for every batch.

comfycomfyuitutorialbatchapiclimcp
Jul 22, 2026·9 min read· Listen

ComfyUI lesson 10: Debug any broken ComfyUI workflow

Start at the first red node and resolve models, nodes, VRAM, shapes, and blank outputs systematically.

comfycomfyuitutorialdebuggingquick fixportability
Jul 21, 2026·7 min read· Listen

Inkling: Thinking Machines’ open-weights model is built to be customized

Inside Inkling’s 975B MoE architecture, controllable reasoning, multimodal training, Apache 2.0 release, Tinker serving limits, and the new direct app.nz gateway route.

modelsthinking-machinesopen-sourcemultimodalreasoninggateway
Jul 21, 2026·11 min read· Listen

Building Polygonalize: stable low-poly video with Go, WASM, Canvas, and Three.js

How we built a free open-source image and video polygonalizer: edge-aware meshes up to 20,000 triangles, custom transparent primitives, stable video topology, Go/WASM, Canvas, Three.js, a CLI, and an app.nz serverless fallback.

open-sourcegolangwasmvideothreejscreative-tools
Jul 20, 2026·6 min read· Listen

Throughput on a shared GPU: fused kernels and a priority scheduler

How we serve chat, image, video, TTS, STT and a LoRA army on shared GPUs — 2× LLM tok/s (fp8+MTP), a fused multi-LoRA kernel (up to 2.4×), a 10.4× admission gate, tier-based VRAM arbitration, and cheaper building blocks (916k embeds/sec, Gemini STT).

performancegpuinferenceschedulingkernels
Jul 20, 2026·6 min read· Listen

Packaging AlayaWorld: an interactive video world model as a self-hosted Cog

Wrapping AlayaLab’s autoregressive world model (LTX-2.3-derived DiT + gemma-3-12b + Depth-Anything-3) as a Cog: camera presets, a warm engine, and an honest look at the LTX-2 Community License.

open-sourcecogvideoworld-modelgpulicensing
Jul 20, 2026·8 min read· Listen

From Artificial Analysis to production: the useful eval loop

How to turn Artificial Analysis, Arena, LiveBench, SWE-bench, and open leaderboards into a runnable shortlist, workload eval, and production model decision.

evalsmodelsbenchmarksartificial-analysisgateway
Jul 20, 2026·7 min read· Listen

Comfy models: from media evals to working GPU inference

Connect image and video benchmark evidence to hosted FLUX, SDXL, LTX, and Wan workflows, reproducible JSON, model provenance, and a global R2 weight cache.

comfyevalsimagevideogpur2
Jul 20, 2026·9 min read· Listen

Anatomy of an AI compute marketplace: tokens, GPU-seconds, and agent-hours on one set of rails

A deep tech dive on app.nz as a compute marketplace: 232 models across 24 providers with price-ordered self-healing failover, GPU placement that learns cold starts and arbitrages clouds under a fixed price, margins enforced as unit tests, and an agent layer that makes the whole market liquid.

infrastructuregpugatewayagentsmarketplacesinvestors
Jul 20, 2026·10 min read· Listen

The next ten years: minds as products, machine-speed markets, and why big labs are the new big tech

A ten-year thesis: AI minds with personality and skills become a product category, robots become software subscriptions, marketplaces go high-frequency and low-fee, humans and governments get augmented, and vertically integrated labs become the new big tech — with sources, and the app.nz strategy behind it.

futurismagentsroboticsmarketplacesinvestorsstrategy
Jul 19, 2026·7 min read· Listen

1ms web search: gobed int8 embeddings, a self-improving index, and agent tools for flash-tier models

How app.nz search answers repeat queries in ~1ms: int8-quantized gobed embeddings at ~600 bytes/doc, a learned index fed by every search and research run, and agent tool design that cheap fast models like DeepSeek actually use well.

searchembeddingsgobedagentsinfrastructure
Jul 18, 2026·6 min read· Listen

Five-minute training runs: capped QLoRA slices and autoresearch on RunPod

Why we cap every training run at 5 minutes: chained resumable QLoRA micro-runs, a GPU-minute budget ledger, a pod reaper, and a karpathy-style hyperparameter autoresearch loop.

trainingqlorarunpodautoresearchinfrastructure
Jul 17, 2026·5 min read· Listen

Designing animation asset galleries for agents

A stable search and download contract for motion, VFX, terrain, 3D objects, REST, CLI, and MCP clients.

agentsmcpanimationapiopen-source
Jul 17, 2026·5 min read· Listen

From reference image to indexed, downloadable 3D object

How the app.nz image-to-3D gallery connects reference images, reconstruction recipes, OBJ downloads, and production generation jobs.

image-to-3d3dobjectsapithreejs
Jul 17, 2026·6 min read· Listen

Opening the app.nz motion, VFX, and terrain commons

Three pose-capture backends, retargetable character motion, configurable Three.js effects, deterministic terrain, and procedural object generators — open, searchable, and available through web, CLI, API, desktop, and MCP.

animationvfx3dthreejsopen-sourcemcp
Jul 16, 2026·4 min read· Listen

Kimi K3: the largest open model, live on the app.nz gateway

Moonshot's 2.8T-parameter Kimi K3 with a flat-priced 1M context window is available now as kimi-k3 — direct Moonshot routing, automatic failover, one metered key.

modelskimimoonshotgatewayopen-source
Jul 15, 2026·7 min read· Listen

polyserve: in-process serverless that packs hundreds of apps into one process

Announcing polyserve, the open-source in-process serverless host behind app.nz: per-second billing via busy_ms and warm_s, hot-swap deploys, scale-to-zero, and autoscaling to Hetzner and RunPod GPUs.

polyserveserverlesshostingopen-source
Jul 15, 2026·8 min read· Listen

Inside the polyserve Go host: fasthttp, unix-socket slots, and zero-drop hot swaps

How one fasthttp process serves many compiled Go apps with microsecond proxy overhead, atomic generation swaps with a 5s drain, and Postgres addons via injected DATABASE_URL.

polyservegolangfasthttphosting
Jul 15, 2026·6 min read· Listen

Serving tiny ONNX models serverlessly, in-process

Why in-process beats a dedicated model server for small ONNX models: warm sessions in one asyncio host, sub-500ms cold starts, honest per-second metering, and RunPod GPU autoscaling for bigger models.

polyserveonnxpythongpu
Jul 15, 2026·5 min read· Listen

polyserve deploy docs: Go, Python, and Bun quickstarts

Per-language deploy quickstarts for polyserve spaces plus the full host control protocol: deploy, usage, health, routing, scale-to-zero, env vars, addons, and the autoscaler contract.

polyservedocshostingserverless
Jul 15, 2026·6 min read· Listen

Interactive world models on app.nz: GPU sessions, ABot-World, and ARDY

World models need a GPU that stays up and talks over a socket. How the new session-cog protocol works, the two open-source world models you can drive today, and the harness that keeps every cog README honest with real runs.

world-modelscogsgpusessionsopen-sourcerunpod
Jul 13, 2026·7 min read· Listen

What is an AI agent cloud?

Why app.nz is broader than a coding agent, model gateway, inference provider, or app host—and how one integrated loop changes the work of shipping AI software.

agentscloudgatewaydeploysinferenceplatform
Jul 7, 2026·4 min read· Listen

Hardening route HTML, desktop deep links, and notebook paths

A small production hardening pass: serving generated SEO HTML, keeping desktop notebook paths inside the workspace, and fixing one-pass deep-link URL encoding.

desktopsecurityseoengineering
Jul 7, 2026·7 min read· Listen

How app.nz hosted training works

The control plane behind app.nz fine-tuning: model catalogues, hardware offers, durable jobs, signed trainer specs, progress callbacks, R2 artifacts, publishing, and deploys.

traininggpurunpodinfrastructure
Jul 7, 2026·6 min read· Listen

Keeping training GPUs busy without wasting memory

The app.nz training worker checklist: conservative VRAM floors, bf16/TF32 defaults, efficient attention, dataloaders, direct artifact I/O, torch.compile tradeoffs, and phase-by-phase memory release.

traininggpuperformancepytorch
Jul 7, 2026·7 min read· Listen

Deploying Rust in under five seconds

Fast Rust deploys come from separating build time from release time: compile once, push a small runtime image or static bundle, then publish by swapping bytes or image refs.

deploysrustcontainersperformance
Jul 7, 2026·7 min read· Listen

How we built isolated containers for agents and CI

The app.nz isolation model: per-run Docker networks, per-step containers, service sidecars, temp Docker auth, label cleanup, timeouts, and remote machines for bigger blast-radius boundaries.

containersagentscisecurity
Jul 7, 2026·8 min read· Listen

How app.nz runs GPUs on demand

Inside the GPU lifecycle: hardware catalogues with VRAM and compute capability, Cog cold starts, local GPU headroom, RunPod pods, schema introspection, metering, and idle reaping.

gpucogsrunpodinference
Jul 7, 2026·7 min read· Listen

Why Docker cold starts are slow

A cold start is scheduling, image pull, Python import, CUDA init, weight loading, compilation, readiness, and first inference. Here is how app.nz reduces each part.

containersgpuperformanceinference
Jul 7, 2026·8 min read· Listen

Dynamic routing across model providers

How the app.nz gateway resolves aliases, classifies auto routes, chooses BYOK or platform keys, walks fallback chains, adapts provider APIs, and meters usage from one catalogue.

gatewaymodelsroutingproviders
Jul 7, 2026·7 min read· Listen

Serverless vs serverful GPU inference

The GPU router behind app.nz cogs and Comfy deployments: local headroom, RunPod serverless, dedicated pods, hysteresis, single-flight cold starts, and when each tier wins.

gpuserverlessroutinginference
Jul 7, 2026·6 min read· Listen

Serving static sites from R2 and a database

The fast path behind app.nz static deploys: upload changed files, prune removed paths, index them in the database, and serve bytes directly from R2 or local storage.

sitesr2deploysperformance
Jul 7, 2026·6 min read· Listen

Build Studio: private container builds for app.nz

How app.nz turns a Dockerfile and build context into a private registry image, with scoped Docker auth, namespaced image refs, redacted logs, and a clean release boundary.

buildscontainersregistrydeploys
Jul 7, 2026·5 min read· Listen

appnz.yaml: the runtime contract

Why app.nz separates static and server runtimes with a small config file instead of guessing forever from package.json, Cargo.toml, dist folders, or Dockerfiles.

deploysconfigruntimebuilds
Jul 7, 2026·6 min read· Listen

CI service containers without shared state

Inside app.nz CI: per-run Docker networks, Postgres/Redis/Mongo service aliases, per-step containers, readiness checks, timeouts, and label-based cleanup.

cicontainerspostgresredis
Jul 7, 2026·6 min read· Listen

How coding agents become durable worker jobs

A coding-agent prompt becomes a stored task, a leased worker job, an event stream, and durable artifacts instead of a long HTTP request that disappears on failure.

agentsworkersqueuesreliability
Jul 7, 2026·6 min read· Listen

Metering the model gateway without guessing

How app.nz meters token streams, image responses, media jobs, BYOK traffic, failures, and routed provider/model pairs from the same catalogue used for pricing.

gatewaybillingmodelspricing
Jul 7, 2026·5 min read· Listen

BYOK provider keys in the model gateway

How app.nz lets users bring provider keys while keeping routing unified, validating providers, preferring user keys, hiding full secrets, and avoiding open-relay behavior.

gatewaybyoksecurityproviders
Jul 7, 2026·5 min read· Listen

Making Cog containers self-describing

How app.nz warms Cog containers, waits for real readiness, reads health and OpenAPI endpoints, and turns model schemas into typed prediction forms.

cogsgpuschemainference
Jul 7, 2026·5 min read· Listen

Mirroring ComfyUI models to make cold starts boring

Why app.nz mirrors popular Comfy model weights into R2 so GPU workers fetch predictable artifacts instead of relying on third-party hosts during cold starts.

comfygpur2performance
Jul 7, 2026·5 min read· Listen

RunPod serverless endpoints under the hood

How app.nz creates scale-to-zero GPU endpoints with bounded workers, queue-delay scaling, endpoint reuse, job polling, image updates, and explicit failure handling.

runpodserverlessgpuinference
Jul 7, 2026·6 min read· Listen

The lifecycle of an app.nz cloud machine

Provisioning is more than create: app.nz tracks provider ids, machine status, ready times, termination, simulated providers, and cleanup for CPU and GPU infrastructure.

provisioninghetznerrunpodinfrastructure
Jul 7, 2026·5 min read· Listen

Metering notebooks by the minute

How app.nz notebooks run free in the browser or on metered CPU/GPU sessions with pricing from the same endpoint, idle stop, balance stop, and durable session rows.

notebooksbillinggpumarimo
Jul 7, 2026·5 min read· Listen

Queues, visibility timeouts, and dead messages

The queue primitive behind app.nz background work: receive as a lease, ack on success, retry after visibility timeout, and move poison messages to dead state.

queuesworkersreliabilityagents
Jul 7, 2026·5 min read· Listen

Durable artifacts for coding agents

Why app.nz stores agent diffs, logs, previews, screenshots, and generated files as artifacts so a run remains auditable after the worker shuts down.

agentsartifactsr2deploys
Jul 13, 2026·4 min read· Listen

Cheaper GPUs by scheduling: multi-cloud arbitrage, hidden cold starts, and cogs that learn their own footprint

How app.nz cog hosting cut serving costs: community-first multi-cloud provisioning under a fixed price, serverless hedging that hides pod cold starts, and per-model learned stats (cold-start EMA, peak VRAM) that feed routing and future GPU bin-packing.

gpueconomicsmulti-cloudserverlesscogsautoscalingomniserve
Jul 10, 2026·4 min read· Listen

Hosted addons: Postgres with pgvector + PostGIS, and gobed GPU vector search

Attach managed services to any app or agent in one command — gobed GPU CAGRA vector search, PostgreSQL with HNSW vectors, PostGIS and graph traversal, Mongo, cache, auth, and AI monitoring — with env vars injected into every runtime.

addonspostgrespgvectorpostgisvector-searchgobedcliagents
Jul 10, 2026·7 min read· Listen

Run cloud agent tasks on your own computers with app worker

Fire an agent task from your phone and let your desktop code it: app worker turns any machine you own into a private, user-scoped worker with local engines — our codex fork (mainline fallback), claude, cursor-agent, gemini, grok.

agentsworkersclilocalcoding
Jul 10, 2026·3 min read· Listen

We open-sourced our investor room

The app.nz data room is now public pages: the 10-year vision, a full technical deep dive (architecture, ~1 ms routing, security, economics, benchmarks, roadmap, risks), all versioned in the repo.

companyinvestorsopen-sourceroutingagents
Jul 10, 2026·5 min read· Listen

Will it fit? An AI agent that 3D bin-packs your life into a van

A chat agent that estimates real item sizes, runs 3D bin-packing algorithms, and renders interactive three.js loading plans — with an honest verdict when it will not fit. Plus a free /api/packing/solve API.

3dpackingagentsthreejsapi
Jul 2, 2026·4 min read· Listen

Hosted notebooks: marimo in the browser or on cloud GPUs, billed per minute

app.nz notebooks run free on Pyodide in your browser or on per-minute CPU/GPU machines that stop themselves when idle. Plus SQLite, Parquet, and agent-trace dataset viewers with HTTP range reads.

notebooksdatasetsgpupythonpricing
Jun 26, 2026·5 min read· Listen

app.nz is now an MCP server

Point Claude, Codex, or any MCP client at app.nz/mcp and your agent can run every gateway model — chat, images, embeddings — plus search the MCP directory, through one connector and one API key.

mcpagentsgatewaytools
Jun 25, 2026·5 min read· Listen

Cloud coding agents: one prompt, one pull request

Run coding agents in sandboxed cloud environments instead of on your laptop — provider-aware (Codex, Claude Code, or our runner), skill-attachable, triggerable by schedules and webhooks.

agentscloudcodingcli
Jun 24, 2026·4 min read· Listen

Instant static deploys: from directory to URL in one command

app.nz hosts static sites behind a URL instantly — app sites deploy uploads, diffs, and prunes a directory so a deploy is the exact state of your build, scriptable into CI or an agent run.

sitesdeployshostingcli
Jun 26, 2026·6 min read· Listen

A searchable directory of 120+ MCP servers

What the Model Context Protocol is, how to wire a server into Claude or Codex, and a new searchable index of 120+ MCP servers filterable by category and pricing.

mcpagentstoolsdirectory
Jun 19, 2026·8 min read· Listen

Test-time training is a linear attention operator in disguise

A small reproduction of a 2026 paper showing that key-value-binding TTT unrolls into a linear-attention state, plus sweeps for extra inner steps, query mismatch, gradient sign, and momentum.

paperstest-time traininglinear attentionsequence models
Jun 19, 2026·7 min read· Listen

An attention sink can mean two opposite things

A small reproduction of a June 2026 paper showing that the same attention stripe can be either a no-op or a broadcast channel.

papersinterpretabilitytransformersattention
Jun 2, 2026·5 min read· Listen

The best coding models in 2026: a hands-on comparison

How to pick a coding model by the job — frontier vs agentic-mid vs fast — measured on cost per merged PR, not cost per token.

modelscodingcomparisonagents
May 26, 2026·4 min read· Listen

Cheap and fast: choosing budget LLMs for high-volume work

A 3-step workflow to find the cheapest model that still clears your quality bar — measured as cost per acceptable output.

modelspricingevalscomparison
May 19, 2026·6 min read· Listen

How to optimize prompts with evals (LLM-as-a-judge)

Treat prompts like code: score them with an eval, then let the optimizer search for a better one against your own dataset.

promptsevalsoptimizationllm-as-judge