Latest
Liquid AI Ships LFM2.5-DSpark, Claims Up to 3.2x Faster Inference
Liquid AI released LFM2.5-DSpark, a model update that runs inference up to 3.2x faster than its prior LFM2.5 generation, according to a Hugging Face blog post from the company. The gain comes from architectural changes Liquid AI describes but has not fully detailed publicly, targeting deployments where response speed and compute cost matter most.
What changes for operators — For a B2B company running an AI chat agent, ticket triage bot, or sales qualification assistant on a small, self-hosted or edge-deployed model, a 3.2x inference speedup translates into lower latency per response and fewer GPU-hours per conversation — meaning either cheaper hosting bills at the same volume, or the ability to run a more capable model at the same cost. Teams currently constrained by response-time SLAs in live chat or voice support, where every second of model latency shows up as customer wait time, get the most immediate benefit; teams using hosted API models from major vendors won't see any change until (or unless) those vendors adopt similar techniques.