KATok Adaptive Tokenizer for Efficient Video Representation
August 24, 2026
KATok is a transformer-based VAE that uses a learned adaptive token selector to improve video compression in latent diffusion models. By discarding uninformative tokens based on content-richness probabilities, it allows for data-dependent compression that adapts to spatio-temporal complexity.
HOW THIS AFFECTS YOU
●
builderThis could significantly reduce the computational cost of training and running video diffusion models.
●
researcherYou can leverage adaptive tokenization to handle varying video complexities more efficiently than fixed-ratio VAEs.