CANIS · Project Page

CANIS: Canonicalize Your 3D Models

Shape-guided canonicalization through semantic 2D–3D correspondence

A semantic, correspondence-driven framework for recovering a consistent canonical orientation from synthetic, real-world, and partial 3D observations.

CANIS teaser showing diverse 3D objects before and after canonicalization
01 / Teaser Arbitrary pose → canonical pose
Scroll to explore
01

Overview · Abstract

A canonical frame for the 3D world.

Canonical orientation is a quiet prerequisite for reliable 3D understanding: objects that mean the same thing should face the same way.

We present CANIS, a framework for canonicalizing 3D objects in arbitrary orientations. CANIS uses a shape-guided generative prior to synthesize a canonical proxy, then extracts semantic 2D–3D correspondences from cross-attention to align the input geometry with that proxy.

The resulting patch-to-cluster registration is category-agnostic and robust across diverse shapes, real-world scans, and partial observations. The same correspondences naturally support interpretable cross-model analysis.

Method at a glance Expand
CANIS pipeline: shape-guided canonical proxy synthesis and semantic correspondence registration
A canonical proxy is synthesized from the input shape. Cross-attention connects selected image patches to semantic 3D clusters for anchored registration.
03

OmniObject3D · Real data

Real-world scans, canonically aligned.

CANIS transfers clean semantic orientation cues to noisy, textured real-world captures. Each group pairs four captured scans with their recovered canonical orientations.

Real dataSet 01
01–04/16
InputCaptured orientation
01
02
03
04
OutputCanonical pose
01
02
03
04
04

Partial observations

Reasoning beyond what is visible.

Even when geometry is incomplete, the semantic proxy provides a stable frame for recovering the intended orientation across groups of challenging partial examples.

Partial dataSet 01
01–04/16
InputPartial geometry
01
02
03
04
OutputCanonical pose
01
02
03
04
06

2D–3D correspondence

Following meaning across representations.

Paired views expose the semantic links learned by different reconstruction backbones. Select a model to inspect its 2D image patches and corresponding 3D regions.

Fish 01 / 10
2D imageSemantic anchors
TRELLIS-OA backbone: Image patch ↔ 3D anchor
3D modelAnchor projection
GLB pending
01

TRELLIS-OA backbone

Image patch ↔ 3D anchor

The OA variant exposes patch-to-cluster bindings used by CANIS to construct semantically anchored 3D registration constraints.

Conclusion

Canonical orientation,
without category boundaries.

Back to top