Eight-bit and four-bit formats are now standard for inference where thirty-two-bit was once assumed. Since inference is usually limited by memory bandwidth rather than arithmetic, halving the bytes per weight nearly halves the time spent waiting.
Quality loss is real but often small, and the tradeoff is usually worth it — which is why quantisation is one of the main levers for running capable models on modest hardware.
