The vLLM Semantic Router team has released Decision 3.0, a family of multimodal decision models. Decision 3.0 models read text, JSON, and images, then answer typed questions about them. There are five sizes, ranging from 0.8B to 27B parameters, all released under Apache-2.0. For developers, this offers a fast way to classify, route, and gate requests without parsing generated text.
TL;DR
- Model sizes: Five variants — d3-lite (0.85B), d3-nano (2.21B), d3-mini (4.54B), d3-flash (8.39B), and d3 (26.09B). Context length: not disclosed.
- Runs on: Latency measured on 1× AMD Instinct MI32…
via MarkTechPost
