Daniel Han’s Post

You can now fine-tune Qwen3.8-27B for free with our notebook! 🔥 Local training works on 24GB VRAM. Unsloth AI trains Qwen3.8 1.5x faster with 50% less VRAM than other setups with FA2. We utilize many kernels including our own and others like Flash Linear Attention kernels for maximum performance. GitHub: https://lnkd.in/gyaDBTxK Qwen3.8-27B Notebooks + Guide: https://lnkd.in/ggPFQgtr

  • graphical user interface, text, application

In case you were wondering, Kaggle, like Colab, provides 30 hours of free GPU with 2× Tesla T4s so you can fine-tune Qwen3.8-27B completely for free by just having a Google account! 🦥 Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.

Shivashish Jaishy

Founder | CEO @ Shristyverse

1mo

24GB is the floor, not the comfortable zone. max_seq_length will blow it before rank will. Keep seq at 2k-4k and r=16 if you actually want a finished run. https://unsloth.ai/docs/models/qwen3.8

fitting qwen3.8-27b training into 24gb is the part worth understanding, not just trusting. is the 50% VRAM cut mostly from the FLA kernels replacing quadratic attention memory, or is gradient checkpointing doing the heavy lifting here. those two produce very different speed/memory tradeoffs at longer sequence lengths.

The 24GB part is impressive, but I’m more curious about the trade-off at longer context lengths. At some point VRAM savings from the kernels can get eaten up by sequence length and activations —would be really interesting to see the throughput/VRAM curve for 4K → 16K → 32K.

This is the kind of optimization that actually matters in production, more capability on practical hardware changes who can build and ship. When 27B class fine tuning fits into a 24GB workflow, the experimentation loop gets a lot more real. What are you seeing on stability and quality once these gains are pushed into real multi turn enterprise workloads?

Like
Reply

Squeezing fine-tuning for a 27B model onto a single 24GB GPU is huge for small research teams and solo developers. When you cut VRAM overhead by 50%, you don't just save money on cloud instances—you radically accelerate the experimental feedback loop for custom domain models. Graet work !

Like
Reply

A 27B model fine-tunable on 24GB VRAM is the part that actually changes who can experiment. Does the 50% VRAM saving hold once you push context toward the 262K limit?

Like
Reply

Unsloth AI always contributing in the open-weights space.

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories