首页 /研究 /VideoGen: Generative Modeling of Videos using VQ-VAE and Transformers
OTHER

VideoGen: Generative Modeling of Videos using VQ-VAE and Transformers

Yunzhi Zhang, Wilson Yan, Pieter Abbeel, Aravind Srinivas

发表年份
2021
引用次数
2

摘要

We present VideoGen: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos. VideoGen uses VQ-VAE that learns learns downsampled discrete latent representations of a video by employing 3D convolutions and axial self-attention. A simple GPT-like architecture is then used to autoregressively model the discrete latents using spatio-temporal position encodings. Despite the simplicity in formulation, ease of training and a light compute requirement, our architecture is able to generate samples competitive with state-of-the-art GAN models for video generation on the BAIR Robot dataset, and generate coherent action-conditioned samples based on experiences gathered from the VizDoom simulator. We hope our proposed architecture serves as a reproducible reference for a minimalistic implementation of transformer based video generation models without requiring industry scale compute resources. Samples are available at https://sites.google.com/view/videogen

关键词

Computer scienceTransformerArchitectureGenerative grammarArtificial intelligenceGenerative modelRobotAction recognitionComputer visionPattern recognition (psychology)

相关论文

查看 OTHER 分类全部论文