Publications
Translational Biomedical AI Controllable Multimodal Generation Multimodal Perception and Understanding
Translational Biomedical AI
CLINES: Clinical LLM-based Information Extraction and Structuring Agent
Translational Biomedical AI Co-first Author
Learning residue-level context for modeling protein-protein interactions
Translational Biomedical AI Co-first Author
A Weakly Supervised Transformer for Rare Disease Diagnosis and Subphenotyping from EHRs with Pulmonary Case Studies
Translational Biomedical AI Co-first Author
Beyond Independent Genes: Learning Module-Inductive Representations for Single-Cell Gene Perturbation Prediction
Translational Biomedical AI Corresponding Author
MedSAM2: Segment Anything in 3D Medical Images and Videos
Translational Biomedical AI Co-first Author
Show full list (10 papers)
CLINES: Clinical LLM-based Information Extraction and Structuring Agent
Translational Biomedical AI Co-first Author
Learning residue-level context for modeling protein-protein interactions
Translational Biomedical AI Co-first Author
A Weakly Supervised Transformer for Rare Disease Diagnosis and Subphenotyping from EHRs with Pulmonary Case Studies
Translational Biomedical AI Co-first Author
Beyond Independent Genes: Learning Module-Inductive Representations for Single-Cell Gene Perturbation Prediction
Translational Biomedical AI Corresponding Author
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
Translational Biomedical AI
Prompt-based multimodal representation learning for drug repurposing
Translational Biomedical AI Corresponding Author
MedSAM2: Segment Anything in 3D Medical Images and Videos
Translational Biomedical AI Co-first Author
X-Field: A Physically Grounded Representation for 3D X-ray Reconstruction
Translational Biomedical AI Spotlight
MuscleParseNet: A Novel Framework for Parsing Muscles of Drosophila Larva in Light-Sheet Fluorescence Microscopy Images
Translational Biomedical AI
Controllable Multimodal Generation
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Controllable Multimodal Generation
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
Controllable Multimodal Generation
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
Controllable Multimodal Generation
Insert Anything: Image Insertion via In-Context Editing in DiT
Controllable Multimodal Generation Oral
Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Controllable Multimodal Generation
3DIS: Depth-Driven Decoupled Image Synthesis for Universal Multi-Instance Generation
Controllable Multimodal Generation Co-first Author Spotlight
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
Controllable Multimodal Generation Corresponding Author
Show full list (32 papers)
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Controllable Multimodal Generation
Photorealistic Text-to-3D Avatar Generation with Constraints for Decoupled Geometry and Appearance
Controllable Multimodal Generation
Toward General-Purpose Video Reconstruction Through Synergy of Grid-Splicing Diffusion and Large Language Models
Controllable Multimodal Generation
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
Controllable Multimodal Generation
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
Controllable Multimodal Generation
GD-NeRF: Generative Detail Compensation for One-shot Generalizable Neural Radiance Fields
Controllable Multimodal Generation Corresponding Author
Test-Time Adaptation for Real-World Video Adverse Weather Restoration With Meta Batch Normalization
Controllable Multimodal Generation
High Fidelity Makeup via 2D and 3D Identity Preservation Net
Controllable Multimodal Generation
Replication in Visual Diffusion Models: A Survey and Outlook
Controllable Multimodal Generation
Insert Anything: Image Insertion via In-Context Editing in DiT
Controllable Multimodal Generation Oral
SKDream: Controllable Multi-view and 3D Generation with Arbitrary Skeletons
Controllable Multimodal Generation Highlight
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
Controllable Multimodal Generation
Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Controllable Multimodal Generation
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
Controllable Multimodal Generation
3DIS: Depth-Driven Decoupled Image Synthesis for Universal Multi-Instance Generation
Controllable Multimodal Generation Co-first Author Spotlight
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
Controllable Multimodal Generation Corresponding Author
Show Me a Video: A Large-Scale Narrated Video Dataset for Coherent Story Illustration
Controllable Multimodal Generation
Controllable 3D Face Generation with Conditional Style Code Diffusion
Controllable Multimodal Generation Corresponding Author
DRIP: Unleashing Diffusion Priors for Joint Foreground and Alpha Prediction in Image Matting
Controllable Multimodal Generation
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
Controllable Multimodal Generation
SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human Reconstruction
Controllable Multimodal Generation Corresponding Author Highlight
Photorealistic Text-to-3D Avatar Generation with Constrained Geometry and Appearance
Controllable Multimodal Generation
Human101: Training 100+ FPS Human Gaussians in 100s from 1 View
Controllable Multimodal Generation
AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D Diffusion
Controllable Multimodal Generation
Global-correlated 3D-decoupling Transformer for Clothed Avatar Reconstruction
Controllable Multimodal Generation
Efficient Emotional Adaptation for Audio-driven Talking-Head Generation
Controllable Multimodal Generation
TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering
Controllable Multimodal Generation
JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh Recovery
Controllable Multimodal Generation
Pyramid Diffusion Models For Low-light Image Enhancement
Controllable Multimodal Generation
Multimodal Perception and Understanding
Efficient training of large vision models via advanced automated progressive learning
Multimodal Perception and Understanding
SELongVLM: Empowering Long Video Language Models with Self-Corrective Clip Selection
Multimodal Perception and Understanding
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
Multimodal Perception and Understanding
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
Multimodal Perception and Understanding Co-first Author
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Multimodal Perception and Understanding First Author
Scalable Video Object Segmentation with Identification Mechanism
Multimodal Perception and Understanding First Author
CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation
Multimodal Perception and Understanding Corresponding Author Best Paper
Show full list (41 papers)
Efficient training of large vision models via advanced automated progressive learning
Multimodal Perception and Understanding
SELongVLM: Empowering Long Video Language Models with Self-Corrective Clip Selection
Multimodal Perception and Understanding
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
Multimodal Perception and Understanding
IDPro: Flexible Interactive Video Object Segmentation by ID-Queried Concurrent Propagation
Multimodal Perception and Understanding
Noise-Tolerant Hybrid Prototypical Learning with Noisy Web Data
Multimodal Perception and Understanding
Exploiting EfficientSAM and Temporal Coherence for Audio-Visual Segmentation
Multimodal Perception and Understanding
The devil is in temporal token: High quality video reasoning segmentation
Multimodal Perception and Understanding
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
Multimodal Perception and Understanding Co-first Author
Few-shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic Segmentation
Multimodal Perception and Understanding
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Multimodal Perception and Understanding First Author
Scalable Video Object Segmentation with Identification Mechanism
Multimodal Perception and Understanding First Author
The First Visual Object Tracking Segmentation VOTS2023 Challenge Results
Multimodal Perception and Understanding
ZJU ReLER Submission for EPIC-KITCHEN Challenge 2023: Semi-Supervised Video Object Segmentation
Multimodal Perception and Understanding
ZJU ReLER Submission for EPIC-KITCHEN Challenge 2023: TREK-150 Single Object Tracking
Multimodal Perception and Understanding
Collaborative Content-Dependent Modeling: A Return to the Roots of Salient Object Detection
Multimodal Perception and Understanding
CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation
Multimodal Perception and Understanding Corresponding Author Best Paper
Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and Segmentation
Multimodal Perception and Understanding
Video Object Segmentation in Panoptic Wild Scenes
Multimodal Perception and Understanding
Co-Learning Meets Stitch-Up for Noisy Multi-Label Visual Recognition
Multimodal Perception and Understanding
FedSeg: Class-Heterogeneous Federated Learning for Semantic Segmentation
Multimodal Perception and Understanding
ProD: Prompting-to-disentangle Domain Knowledge for Cross-domain Few-shot Image Classification
Multimodal Perception and Understanding
Decompose to Generalize: Species-Generalized Animal Pose Estimation
Multimodal Perception and Understanding
The Tenth Visual Object Tracking VOT2022 Challenge Results
Multimodal Perception and Understanding
V²L: Leveraging Vision and Vision-language Models into Large-scale Product Retrieval
Multimodal Perception and Understanding
Decoupling Features in Hierarchical Propagation for Video Object Segmentation
Multimodal Perception and Understanding First Author Spotlight
Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation
Multimodal Perception and Understanding
In-N-Out Generative Learning for Dense Unsupervised Video Segmentation
Multimodal Perception and Understanding
H2FA R-CNN: Holistic and Hierarchical Feature Alignment for Cross-domain Weakly Supervised Object Detection
Multimodal Perception and Understanding
Rethinking cross-modal interaction from a top-down perspective for referring video object segmentation
Multimodal Perception and Understanding
Towards multi-object association from foreground-background integration
Multimodal Perception and Understanding First Author
Associating Objects with Transformers for Video Object Segmentation
Multimodal Perception and Understanding First Author
Collaborative Video Object Segmentation by Multi-Scale Foreground-Background Integration
Multimodal Perception and Understanding First Author
DSC-PoseNet: Learning 6DoF Object Pose Estimation via Dual-scale Consistency
Multimodal Perception and Understanding First Author
Memory aggregated cfbi+ for interactive video object segmentation
Multimodal Perception and Understanding
CFBI+: Collaborative video object segmentation by multi-scale foreground-background integration
Multimodal Perception and Understanding First Author
Collaborative Video Object Segmentation by Foreground-Background Integration
Multimodal Perception and Understanding First Author Spotlight
Gated Channel Transformation for Visual Recognition
Multimodal Perception and Understanding First Author
Going Deeper Into Embedding Learning for Video Object Segmentation
Multimodal Perception and Understanding First Author
Dual Embedding Learning for Video Instance Segmentation
Multimodal Perception and Understanding
