Skip to project
Hwan Heo / Selected work & writing
Projects

01 / PROJECT CASE STUDY

VARCO 3D:Production 3D Generative AI

Project Overview

VARCO 3D is NC AI’s production 3D generative AI service for creating detailed meshes with 4K PBR-aware textures, AI auto-rigging and retargeting, and decimation with normals baked from native high-poly detail.

I led the overall research direction while working across the full generation stack: training in-house 3D generative models with more than 1B parameters, optimizing their production inference paths, and building GPU-native mesh and texture-processing systems for asset delivery.

Ten recent creations published on VARCO 3D Explore.

Core contributions

  • Led the overall VARCO 3D research direction and trained dense and sparse in-house 3D generative models with more than 1B parameters.
  • Reduced 15-step denoising latency by 25.66% without quantization through forward-path analysis and CUDA kernel optimization.
  • Built a CUDA mesh pipeline that cleans, remeshes, and decimates geometry from 1M to 1K faces in approximately five seconds.
  • Developed a tile-based UV unwrapper that processes meshes with more than 1M faces in under 30 seconds on average, while maintaining over 50% UV-space utilization even on challenging inputs.
  • Implemented a rasterization-based CUDA multi-view texture-projection pipeline that generates an 8K texture map in approximately two seconds.

System Focus

The practical target of VARCO 3D is not just to synthesize geometry. The service path extends native high-poly output into 4K PBR-aware textured assets, detail-preserving decimation with baked normals, and animation workflows through AI auto-rigging and retargeting.

I worked across three connected layers, from setting the research direction and training 1B+ parameter models from scratch to shipping the GPU-native systems around them:

  1. Native 3D generation — train large in-house geometry models rather than relying on slow per-asset SDS optimization.
  2. Production inference — profile and rewrite the expensive denoising path so the model can run under service latency constraints.
  3. Mesh and texture processing — convert generated geometry into cleaned, decimated, unwrapped, textured mesh assets.

Architecture Shift

The model stack moved from dense latent generation to sparse active-structure generation. That shift mattered because it changed both the training target and the serving bottleneck.

VARCO 3D 1.0

VecSet-based ShapeVAE and Dense DiT Denoiser

The first stack used dense geometry generation with lattice-conditioned refinement, focused on producing stable mesh structure from native 3D latents.

VARCO 3D 2.0

Sparse DC VAE and Sparse DiT Denoiser

The second stack moved generation onto active sparse 3D structure for higher-detail outputs and a more scalable inference path. Sparse assets are less uniform at runtime, however: token count and active layout vary by input. Serving therefore needed model-aware profiling and custom CUDA work rather than only generic graph-level acceleration.

Production Inference Optimization

On the VARCO 3D 2.0 sparse denoiser, I profiled the forward path and identified null-context attention as arithmetic the model did not need to repeat. The optimized path replaces unconditional cross-attention with a fixed-vector path, then fuses memory-bound tensor operations through custom CUDA kernels and cuBLASLt epilogues.

Across ten production assets with 3,664 to 30,227 active tokens, the combined path reduced CUDA-synchronized 15-step denoising latency by 25.66% on average on A100 BF16. The gain came from forward-path and kernel optimization alone, without quantizing the model.

This was not a benchmark-only shortcut. The optimized path entered production serving with numerical validation and fallback rules around the fused kernels.

Read the VARCO 3D 2.0 inference optimization write-up

Mesh and Texture Delivery

To extend geometry generation into textured-mesh delivery, I implemented and optimized the post-generation pipeline around GPU execution:

  • CUDA-based topology cleaning, remeshing, and QEM decimation that robustly reduces meshes from 1M to 1K faces in about five seconds.
  • Custom UDF kernels for fast solid correction from generated geometry.
  • Flood-fill acceleration kernels to make occupancy and solidness correction practical at service scale.
  • Optimized tile-based UV unwrapping that processes meshes with over 1M faces in under 30 seconds on average while keeping UV-space utilization above 50% even in worst-case inputs.
  • CUDA rasterizer–based back-projection with visibility-aware view selection and blending that projects multi-view generated images into an 8K UV texture map in approximately two seconds.

This layer is where the generated result becomes a deployable asset: cleaned geometry, controlled face count, valid UVs, and texture maps produced without a slow CPU-bound handoff.