CVPR 2026

HAMMER

Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

Lei Yao1, Yong Chen2, Yuejiao Su1, Yi Wang1, Moyun Liu2, Lap-Pui Chau1

1The Hong Kong Polytechnic University  ·  2Huazhong University of Science and Technology

+5.4
aIOU · PIAD unseen
+9.1
AUC · PIAD unseen
7/7
corruptions · best
LOADING…
drag to orbit · space to flip
One photo, one region. The interaction thumbnail is all the model sees of the intention — no attribute description, no 2D segmenter. Flip between the prediction and the annotation and judge it yourself.
TL;DR

HAMMER reads the interaction intention out of a single image into a contact-aware embedding, has an MLLM name the affordance, and folds that knowledge into the point cloud through hierarchical cross-modal integration and multi-granular geometry lifting.

01

Abstract

Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate for a novel framework that leverages emerging multimodal large language models (MLLMs) for interaction intention-driven 3D affordance grounding, namely HAMMER.

Instead of generating explicit object attribute descriptions or relying on off-the-shelf 2D segmenters, we alternatively aggregate the interaction intention depicted in the image into a contact-aware embedding and guide the model to infer textual affordance labels, ensuring it thoroughly excavates object semantics and contextual cues.

We further devise a hierarchical cross-modal integration mechanism to fully exploit the complementary information from the MLLM for 3D representation refinement, and introduce a multi-granular geometry lifting module that infuses spatial characteristics into the extracted intention embedding, thus facilitating accurate 3D affordance localization.

Extensive experiments on public datasets and our newly constructed corrupted benchmark demonstrate the superiority and robustness of HAMMER compared to existing approaches. All code and weights are publicly available.

02

How it works

HAMMER architecture: an MLLM turns the interaction image
             and a prompt into a contact-aware embedding, which is folded into the point cloud by
             hierarchical cross-modal integration and multi-granular geometry lifting before the
             affordance decoder.
Architecture. The interaction image and an object-centric prompt go to an MLLM; the [CONT] token collects the interaction into 𝐟c while the model also names the affordance in words. That knowledge enters the point cloud through hierarchical cross-modal integration, and the embedding is given geometry back by multi-granular lifting before the affordance decoder produces the map.
01 · Intention embedding

A vocabulary token [CONT] is added purely to collect interaction-related information; its hidden state becomes the contact-aware embedding 𝐟c. Naming the affordance in words is kept as an auxiliary task, which is what forces the token to carry the interaction.

02 · Hierarchical integration

MLLM hidden states meet the point cloud twice — once at the encoder bottleneck through cross-attention, and again at full resolution through a gated global descriptor. One fusion at the end would have skipped the contextual half.

03 · Geometry lifting

𝐟c came from an image, so it knows the intention but not the shape. It walks the decoder's scales from coarse to fine, querying point features at each one. No camera parameters and no intermediate 2D contact map.

04 · Affordance decoding

Enhanced point features attend to the 3D-aware embedding, and an MLP with a sigmoid turns the result into a score per point. Training adds the language loss to a focal-plus-dice affordance loss.

03

Results

Fig. 01 · PIAD — seen categories

LOADING…
every panel orbits on its own · space to flip
Seen categories. The thumbnail in each panel is the interaction the model was given. Flipping to the annotation shows how much of the difference is real.

Fig. 02 · PIAD — unseen categories

LOADING…
object categories never seen in training
Unseen categories. These object classes are absent from training; the affordance still lands on the right region.

Fig. 03 · PIAD-v2

LOADING…
a larger vocabulary of objects and affordances
PIAD-v2. More categories, more affordances, and a much longer tail.

Fig. 04 · the feature field behind the prediction

LOADING…
per-point features, projected to three principal components
What the points know. These are the learned per-point features, reduced to three dimensions and read as colour. The hues carry no meaning on their own — what matters is that parts of an object settle into their own blocks of colour, which is the structure the affordance head reads off. Flip to affordance to see what it reads out of them.

Fig. 05 · one chair, seven corruptions

LOADING…
all seven panels share one camera
Corrupted PIAD. One chair, asked the same thing — where do you take hold to move it — after being jittered, rotated, squashed, thinned and padded with stray points. Drag the severity up and watch whether the backrest stays lit. Scale looks like a squash rather than a resize because the benchmark renormalises each corrupted cloud to a unit sphere. The numbers are in the table below.

PIAD

aIOU, AUC and SIM higher is better; MAE lower is better.

SeenUnseen
MethodVenue aIOU ↑AUC ↑SIM ↑MAE ↓ aIOU ↑AUC ↑SIM ↑MAE ↓
Intention-driven 3D affordance grounding
PMFICCV 2110.1375.050.4250.1414.6760.250.3300.211
ILNSIGGRAPH 2211.2575.840.4270.1374.7159.690.3250.207
FRCNNRemote Sensing 2311.9776.050.4290.1364.7161.920.3320.195
PFusionCVPR 1812.3177.500.4320.1355.3361.870.3300.193
XMFNetNeurIPS 2212.9478.250.4410.1275.6862.580.3420.188
IAGNetICCV 2320.5184.850.5450.0987.9571.840.3520.127
GREATCVPR 2519.6185.220.5690.0938.3267.460.3300.121
HAMMER22.2088.430.6050.08313.7180.920.4490.109
Δ vs. best prior+1.69+3.21+0.036−0.010+5.39+9.06+0.097−0.012
Language-driven 3D affordance grounding · shown for reference
LASOCVPR 2419.7084.200.5900.0968.0069.200.3860.118
GEALCVPR 2522.5085.000.6000.0928.7072.500.3900.102

PIAD-v2

Held out by object and by affordance, separately.

SeenUnseen objectUnseen affordance
MethodYear aIOU ↑AUC ↑SIM ↑MAE ↓ aIOU ↑AUC ↑SIM ↑MAE ↓ aIOU ↑AUC ↑SIM ↑MAE ↓
OpenAD202331.8889.540.5260.10416.6273.490.3390.1598.0061.220.2290.167
FRCNN202333.5587.050.6000.08218.0872.200.3620.1527.9659.080.2100.156
XMFNet202233.9187.390.6040.07817.4074.610.3610.1268.1160.990.2250.152
IAGNet202434.2989.030.6230.07616.7873.030.3510.1238.9962.290.2510.141
LASO202434.8890.340.6270.07716.0573.320.3540.1238.3764.070.2280.140
GREAT202537.6191.240.6600.06719.1676.900.3840.11912.7869.260.2970.135
HAMMER40.0694.190.6980.06324.2884.780.4490.11213.2872.210.3020.128

Robustness

Corrupted PIAD, against the strongest prior method.

aIOU ↑AUC ↑SIM ↑MAE ↓
CorruptionGREATOursGREATOursGREATOursGREATOurs
Scale13.4119.1877.7785.730.4900.5960.1160.093
Jitter11.6717.3675.6584.570.4610.5680.1150.095
Rotate12.6118.3376.9985.310.4830.5910.1160.095
Drop-local8.6717.9872.3884.590.4200.5680.1220.098
Drop-global13.6220.3878.2686.760.4940.6110.1120.089
Add-local11.5017.6775.1686.040.4380.5520.1090.093
Add-global9.6319.0971.2586.360.4320.5780.1080.090
04

Citation

@inproceedings{yao2026hammer,
  title={HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding},
  author={Yao, Lei and Chen, Yong and Su, Yuejiao and Wang, Yi and Liu, Moyun and Chau, Lap-Pui},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2026}
}