Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Framework responds to complaints that BIOS update bricks Ryzen 7040 laptops

    Amazon aims for delivery drones to reach 500 US neighborhoods by end of 2026

    Watch Valve set up the Steam Frame in its own leaked videos

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»Free AI Tools»LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
    Free AI Tools

    LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

    By No Comments3 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Leonie Monigatti's avatar


    Today, we release QAD Q4_0 GGUFs. These are updated 4-bit checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. They allow developers to run LFM2.5 models at Q4_0 memory and speed without the usual quality drop:

    • Trained with Quantization-Aware Distillation (QAD): a high-precision teacher model is distilled into a quantized student model
    • Same memory and speed as native Q4_0: They keep the low memory footprint and high throughput of Q4_0 GGUFs
    • Recovery: 97% of their BF16 average accuracy lost to quantization is recovered



    Benchmark results

    For all four models, we compare their released GGUFs produced with post-training quantization (PTQ) against the trained QAD Q4_0 checkpoints on a benchmark suite spanning reasoning, instruction-following, tool use, and agentic capabilities: GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4. The BF16 GGUF serves as the in-format ceiling. We also add one scale-appropriate math evaluation: GSM8K for LFM2.5-230M and LFM2.5-350M, and AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B. We report the mean across five repeats.

    lfm25_gguf_eval_scores_2x2_liquid

    Across all four models, QAD substantially improves the Q4_0 checkpoint. The QAD checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of their respective BF16 baseline performance.



    Speed and size on real edge hardware

    We measure decode throughput for the four models LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B across four targets: MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5. MacBook Pro and NucBox use GPU inference, while Samsung and Raspberry Pi use Arm CPU inference. BF16 and F16 are shown as full-precision references where profiled.

    lfm25_230m_decode_throughput_4panel_liquid

    lfm25_350m_decode_throughput_4panel_liquid

    lfm25_1p2b_decode_throughput_4panel_liquid

    lfm25_2p6b_decode_throughput_3panel_liquid

    The 230M and 350M QAD Q4_0 checkpoints match Q5_K_M quality within evaluation variance at a 4-33% higher decode throughput. The 1.2B and 2.6B QAD Q4_0 checkpoints match Q4_K_M quality at a 3-14% higher throughput. The QAD Q4_0 checkpoints also match Unsloth’s UD-Q4_K_XL (where applicable, for the 230M and 1.2B), a strong external post-training quantization checkpoint.



    How to use QAD GGUFs

    Use the files with llama.cpp or any runtime that supports GGUF Q4_0 artifacts.

    llama-cli -hf LiquidAI/LFM2.5-350M 
      --hf-file LFM2.5-350M-QAD-Q4_0.gguf 
      -p "What is C. elegans?"
    



    Get Started with QAD GGUFs

    The QAD GGUFs are available on Hugging Face today: LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.

    We can’t wait to see what you build.



    Citation

    For citations, please use the following reference or BibTeX:

    Liquid AI, "LFM2.5 Q4_0: Quantization-Aware Distillation for Edge Deployment", Liquid AI Blog, Aug 2026.
    

    Or use the BibTeX citation

    @article{liquidAI2026Q40,
      author = {Liquid AI},
      title = {LFM2.5 Q4_0: Quantization-Aware Distillation for Edge Deployment},
      journal = {Liquid AI Blog},
      year = {2026},
      note = {www.liquid.ai/blog/qad},
    }
    

    Checkpoints Distillation LFM2.5 Q4_0 QuantizationAware
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleThe women’s soccer league trying to fix fantasy sports
    Next Article Trump expected to pick conservative policy wonk Heidi Overton to lead FDA
    • Website

    Related Posts

    Free AI Tools

    I Saw the Future of AI in a Robot That Can Learn on the Spot

    Free AI Tools

    Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks

    Free AI Tools

    Pacing comes to the AI frontier

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Framework responds to complaints that BIOS update bricks Ryzen 7040 laptops

    0 Views

    Amazon aims for delivery drones to reach 500 US neighborhoods by end of 2026

    0 Views

    Watch Valve set up the Steam Frame in its own leaked videos

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Framework responds to complaints that BIOS update bricks Ryzen 7040 laptops

    0 Views

    Amazon aims for delivery drones to reach 500 US neighborhoods by end of 2026

    0 Views

    Watch Valve set up the Steam Frame in its own leaked videos

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.