app.nzapp
AppsProjectsReposPullsChatIntegrationsGatewayModelsEvalsToolsDatasetsMCPDeploysPricingBlogDocsAssistantsCharactersArtMusic
Sign inStart building
Agent stack
Cloud coding agentAgents SDKIntegrationsBrowser agentMonitors & auto-agentsSchedulersAgent skillsMCP serversDeep research
Models & API
AI GatewayModel catalogModel evalsModel spacesPlaygroundText to imageImage to 3DText to 3DMusic & SFXAudio editorMedia optimizerAI art & libraryChatAPI referenceSchemaBecome a provider
Compute & hosting
DeploysAddonsPostgres hostinggobed vector searchSite hostingAnalyticsCog GPU hostingRL trainingBuilds & CIWorkersTask queuesDomainsGit hosting
Tools
AI toolsDrawDiffusion canvasLive DrawWriteSheetsArtifactsVideo studioNotebooksDatasets
Learn
DocsBlogEval guidesPrompt libraryCLIAlternativesPapersAI charactersArt gallerySecurityConsulting
Company
PricingEnterpriseSettingsBillingStatusInvestorsCreate accountTerms of ServicePrivacy Policy
app.nzapp.nz

The agent cloud for shipping software on app.nz.

Built in New Zealand by App AI NZ.

Social
X / TwitterGitHubYouTube
The app.nz network
GpuBrainPapersReading TimemojojojoNetwrckText-Generator.ioCodex InfinityOpenPathsCuteDSLAI Art GeneratorAIArt-Generator.artSiteSimSimplexGenDictatorFlowWebFiddleRing.nzChatGibidyBitBankExperimentFlowEvangelerHires.nzHow.nzV5 GamesAddicting Word GamesBig Multiplayer ChessWord SmashingreWord GameMultiplication Master
© 2026 App AI NZ Ltd. All rights reserved.All systems normalTermsPrivacy
Blog
August 20, 2026·10 min read·app.nz

A GPU-native video depth and normal-map Cog: NVDEC, DA3, and NVENC

How we built an open video-to-depth Cog with a CUDA-only raw-pixel path, stable clip-wide depth, normal maps, bounded GPU memory, immutable weights, and scale-to-zero deployment.

Listen to this article

On-device voice

Uses the voice built into your browser; no article text leaves this page.

Audio narration is not supported by this browser.

We wanted a small open video model that demonstrates a real GPU media path, not another image demo wrapped in a frame loop. The result is `appnz-video-depth-vfx`: a portable Cog that turns a short video into a stable depth pass, heatmap, tangent-space normal map, or original/depth split preview.

The interesting part is the boundary between the codec and the model:

compressed input
    → NVDEC
    → CUDA uint8 frames
    → Depth Anything 3
    → CUDA depth / heatmap / normal math
    → NVENC
    → compressed video
    → audio remux

Raw pixels stay on the GPU. FFmpeg appears only at the end, where it copies the already-compressed video and optional source audio into the final MP4.

Why Depth Anything 3, not SAM 2

SAM 2 is a strong primitive when the output is an object mask that must remain tracked through time. A normal map needs something different: a continuous surface field. We therefore use ByteDance's Depth Anything 3, specifically the Apache-2.0 `DA3MONO-LARGE` checkpoint.

That choice is also a license boundary. The Mono checkpoint and upstream code are Apache-2.0. DA3's main any-view weights carry noncommercial terms, so this Cog intentionally does not download or expose them.

The checkpoint is pinned by repository revision and baked into a separate Cog weight layer. Cold workers do not contact Hugging Face, and a rebuild cannot silently pick up a different model.

Loading only the inference network

DA3's high-level API imports its exporters and reconstruction utilities. Those paths are useful for GLB, COLMAP, Gaussian splats, and visualisation, but they pull packages such as Open3D, pycolmap, moviepy, and pose tooling into a Cog that only needs monocular depth.

The runner reads the checkpoint's config.json, constructs DepthAnything3Net through DA3's own config factory, strips the wrapper's exact model. safetensors prefix, and loads with a strict missing/unexpected key check. This keeps the official architecture and weights while avoiding the unused export dependency tree.

We verified the pinned safetensors header before relying on that loader: 406 tensors, all under the wrapper prefix. The built container reconstructs a 334,171,394-parameter DepthAnything3Net with no missing or unexpected keys. A mismatch fails setup instead of partially loading a model and producing plausible-looking nonsense.

NVDEC without the accidental CPU path

TorchCodec 0.10 can decode into CUDA tensors and encode CUDA tensors through NVIDIA's hardware codecs. It can also fall back to CPU when NVDEC or the source codec is unavailable. That is a good compatibility default for general applications and a bad hidden surprise for a GPU-path example.

The CUDA 12.8 wheel also links against NVIDIA Performance Primitives. The image pins the NPP 12 runtime explicitly and registers NVIDIA's Python-wheel library directories with the dynamic linker. Importing TorchCodec inside the final container is a release gate, because a host-only unit suite cannot reveal an undiscoverable NPP shared library.

The version matrix is part of the deployment contract, not incidental build metadata. CUDA 13 with PyTorch 2.11 imported on our development GPU but the hosted driver exposed no CUDA device. PyTorch 2.8 with TorchCodec 0.7 restored host compatibility but predates TorchCodec's VideoEncoder. An unqualified TorchCodec 0.10 wheel supplied the encoder but not CUDA decoding. The final image therefore pins the hashed official +cu128 wheels for PyTorch 2.10, TorchVision 0.25, and TorchCodec 0.10, plus NPP 12. That exact set provides both NVDEC and the tensor-native encoder on the app.nz worker fleet.

Two typed codec details were worth testing in the built image as well. FFmpeg's rc and preset values cross a C++ enum boundary in TorchCodec 0.10, so the runner passes numeric VBR and P4 values instead of the familiar "vbr" and "p4" strings. The strings look natural in Python but are rejected before NVENC opens.

The Cog therefore checks both signals after each decode batch:

  • decoder.cpu_fallback must be false when strict_gpu=true;
  • the returned frame tensor must report device.type == "cuda".

4K video is not the problem; hundreds of uncompressed 4K frames at once are. Decode happens in bounded 32-frame chunks and each chunk is resized on CUDA before joining the output-sized frame buffer. This avoids a 12–15GB temporary allocation for a 20-second 4K clip.

Turning a frame sequence into stable depth

DA3 accepts (B, N, 3, H, W), so inference sends short frame groups rather than pretending every frame is an unrelated photograph. Frames are resized to a patch-14 grid, normalised with ImageNet statistics, and inferred under BF16 autocast where the GPU supports it, otherwise FP16.

Raw monocular depth has an arbitrary scale. Normalising every frame separately makes video brightness pump even when the camera is still. We sample the whole clip, take shared 2nd and 98th percentile bounds, and apply one scale to every frame. A small causal EMA then reduces frame-to-frame shimmer:

smoothed[t] = amount × smoothed[t-1] + (1 - amount) × depth[t]

The default amount is 0.15. Higher values stabilise locked-off shots but can trail fast motion, so it remains an input rather than a hard-coded aesthetic.

Normals come directly from finite differences in the normalised depth field:

n = normalize((-strength × dx, -strength × dy, 1))
rgb = 0.5 × n + 0.5

Depth grayscale, the piecewise Turbo-style heatmap, normal conversion, split composition, resize, and uint8 conversion are all Torch operations on CUDA.

Memory is part of the API contract

TorchCodec's simple VideoEncoder currently accepts the complete NCHW tensor. That makes the example compact, but it means output frames and rendered frames coexist before NVENC starts. The contract therefore limits inputs to 20 seconds, 600 source frames, and 4K, then applies a second 7GiB pixel-buffer gate after FPS sampling and output-size selection.

A long 4K source can still produce a 1080p result. Asking to preserve all 4K frames on a 16GB worker fails early with an instruction to lower resolution, FPS, or duration. An explicit refusal is better than a CUDA OOM after the model has already done most of the work.

Cog, GHCR, and app.nz scale to zero

The project uses Cog's current BaseRunner/run interface and pins Python, PyTorch, TorchCodec, DA3 source, and weights. Twelve weight-free tests cover frame sampling, patch geometry, global normalisation, temporal smoothing, normal direction, render modes, video bounds, the output memory budget, and deployment/schema consistency.

Every main-branch or version-tag build publishes an immutable commit image to GHCR, with the model checkpoint in a separate weight layer. The repository also ships an app.nz schema for the video picker, pass selector, sliders, codec controls, and video output.

payload="$(jq '{name,image,hardware,idleSeconds,minVramGb,schema:{outputKind:.schema.outputKind,inputs:[.schema.inputs[]|{name,type,description,default,required,choices,min,max,order}|with_entries(select(.value != null))]}}' deployment.json)"
app api POST /api/cogs "$payload"
app cogs predict MODEL_ID --input '{
  "video":"https://example.com/clip.mp4",
  "effect":"normals",
  "process_resolution":504,
  "target_fps":24
}'
app cogs sleep MODEL_ID

The registered model has a 16GB VRAM floor and a three-minute idle window. The first request boots a GPU worker; sleep or the idle reaper returns it to zero.

The cheaper agent route that built it

This work also changed the app.nz coding agent's no-plan default to DeepSeek V4 Flash through OpenRouter's :floor variant. OpenRouter documents :floor as sorting healthy providers by price; unlike its default tool-routing behaviour, the suffix keeps price ordering explicit for tool calls too.

The app.nz catalogue names that route or/deepseek-v4-flash-floor; a direct OPENROUTER_API_KEY uses the upstream deepseek/deepseek-v4-flash:floor slug. New cloud tasks default to the model's low reasoning tier, while synced ChatGPT plans still use their quota-backed Codex route. Provider-specific failures advance cleanly instead of sending an app.nz-only fallback alias to OpenRouter or OpenAI.

Cheap routing is not a substitute for the engineering gates above. It makes agent loops affordable; pinned sources, bounded memory, device assertions, contract tests, and a real GPU smoke are what make the generated system worth deploying.

Read the source and deployment guide, or open Cog Studio to deploy the published image.

Build what you just read

Ship agents, models, and apps on one cloud.

Start with free credits, then use the same platform from the web app, CLI, desktop app, or MCP.

Start building freeRead the docs

Keep reading

Packaging MiniMax H3 for Cog, R2, RTX 5090, and scale to zero

A careful H3 video workflow with text, keyframes, native audio, true loops, verified R2 weights, GPU AV1, serverless pricing, and a fail-closed license gate.

Packaging AlayaWorld: an interactive video world model as a self-hosted Cog

Wrapping AlayaLab’s autoregressive world model (LTX-2.3-derived DiT + gemma-3-12b + Depth-Anything-3) as a Cog: camera presets, a warm engine, and an honest look at the LTX-2 Community License.

From ComfyUI to any Cog: open audio workflows that scale to zero

How template:name connects ComfyUI to reusable Cog deployments, when to choose open Pocket TTS versus fal audio, and how the shared serverless/pod router returns idle workers to zero.