AI News a compressed, 4-bit model that outperforms its full-precision original Making a large language model smaller almost always comes with a cost. The now-standard recipe for efficient deployment is to…