Skip to content

Infrastructure & pricing, read from an operations desk

Everything we have published under Infrastructure & pricing, read from an operations desk: what it changes for a B2B company of 10-200 people.

  1. Latest

    Liquid AI Ships LFM2.5-DSpark, Claims Up to 3.2x Faster Inference

    Liquid AI released LFM2.5-DSpark, a model update that runs inference up to 3.2x faster than its prior LFM2.5 generation, according to a Hugging Face blog post from the company. The gain comes from architectural changes Liquid AI describes but has not fully detailed publicly, targeting deployments where response speed and compute cost matter most.

    What changes for operatorsFor a B2B company running an AI chat agent, ticket triage bot, or sales qualification assistant on a small, self-hosted or edge-deployed model, a 3.2x inference speedup translates into lower latency per response and fewer GPU-hours per conversation — meaning either cheaper hosting bills at the same volume, or the ability to run a more capable model at the same cost. Teams currently constrained by response-time SLAs in live chat or voice support, where every second of model latency shows up as customer wait time, get the most immediate benefit; teams using hosted API models from major vendors won't see any change until (or unless) those vendors adopt similar techniques.

  1. Researchers Squeeze 33 Points of GPU Utilization Out of Existing Hardware — By Reordering Jobs

    Most 10-200 person B2B companies don't run their own GPU clusters, so this isn't a direct action item — but it's a useful data point when a vendor tells you that scaling an AI feature requires a costly infrastructure upgrade. If job scheduling alone can unlock 33 points of utilization on the same hardware, ask any provider quoting you for "more compute" whether they've actually optimized what they have first. The angle here is procurement leverage and vendor scrutiny, not internal ops change — most readers won't touch a scheduler themselves, but they will pay for one indirectly through inference or fine-tuning costs.

Next step

Free AI Diagnostic

Fifteen minutes, no email required. It maps where your work actually goes and ranks what is worth automating first.

Start the free diagnostic

Starts immediately in the browser.

Fee
Free
Length
15 minutes

You keep the ranked list of candidates either way.