English
← Back to AI Technology

Model architecture

Mixture of Experts (MoE)

A family of routed expert architectures. Modern sparse variants activate only selected expert subnetworks for each input, increasing capacity without activating every parameter on every token.

Reference period
1991 / 2017 / 2021
Last reviewed

Key terms

  • experts
  • router
  • sparse activation
  • load balancing

Connected across the project

Primary sources

  1. Adaptive Mixtures of Local Experts
  2. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
  3. Switch Transformers

This is a curated technical map, not a claim of comprehensive coverage.