DeepSpeech

DeepSpeech is an open-source speech-to-text engine developed by Mozilla that uses deep learning techniques to convert spoken language into written text. It is designed to enable efficient and accurate automatic speech recognition (ASR) accessible to developers and researchers.

Read More →

Latent diffusion model (LDM)

A latent diffusion model (LDM) is a type of generative machine learning model that performs diffusion processes in a compressed latent space, enabling efficient and high-quality image synthesis and related tasks. By operating in a lower-dimensional representation, LDMs reduce computational costs while maintaining detailed output.

Read More →