Music & Audio3 min reading time

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

Apple Research Blog
Read full post
Apple's Siri Expressive Voices use a memory-efficient audio synthesis model called AFM 3 Core Advanced, running on the Apple Matrix Coprocessor. The model employs a novel detokenizer with decoupled temporal and depth processing, enabling real-time, high-fidelity speech synthesis with low memory usage. This architecture improves voice quality and supports customizable voice features on Apple devices.

More in Music & Audio

Music & Audio2 min read

Universal Music is launching an AI music platform with ElevenLabs

Covered by 3 sources
Music & Audio4 min read

Suno trained its v6 AI music models with help from Warner and BMG

Covered by 5 sources
Music & Audio3 min read

Yoto just announced two new audio devices for kids

Engadget