Rethinking cross-modal interaction from a top-down perspective for referring video object segmentation