Adobe Research uses state-space models to handle long-range dependencies, adds dense local attention to maintain coherence, and uses training strategies such as diffusion forcing and frame-local attention to address the long-term memory problem in video generation. They call this a long-standing challenge.