Public training and inference code for a single-GPU text-to-sign-language diffusion model, with a Hugging Face checkpoint.
Kernel-level work on MoDiff: fused low-bit operators and cache-update fusion for measured, not only counted, diffusion speedups.