Toward General-Purpose Video Reconstruction Through Synergy of Grid-Splicing Diffusion and Large Language Models