Exploiting EfficientSAM and Temporal Coherence for Audio-Visual SegmentationShare on Twitter Facebook LinkedIn Previous Next