These startups are chasing the next big thing in LLMs

MIT Technology Review
Read full post
Since the introduction of transformer neural networks in 2017, they have powered all major large language models but now face limitations in handling long text efficiently. Startups are exploring new architectures beyond transformers to improve LLM capabilities and reduce computational costs. This shift aims to address the high energy consumption and scaling challenges inherent in current transformer-based models.

More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 3 sources

Harvey raises $550M more to develop AI tools for legal teams

SiliconANGLE