You can now fine-tune Qwen3.8-27B for free with our notebook! 🔥 Local training works on 24GB VRAM. Unsloth AI trains Qwen3.8 1.5x faster with 50% less VRAM than other setups with FA2. We utilize many kernels including our own and others like Flash Linear Attention kernels for maximum performance. GitHub: https://lnkd.in/gyaDBTxK Qwen3.8-27B Notebooks + Guide: https://lnkd.in/ggPFQgtr
24GB is the floor, not the comfortable zone. max_seq_length will blow it before rank will. Keep seq at 2k-4k and r=16 if you actually want a finished run. https://unsloth.ai/docs/models/qwen3.8
fitting qwen3.8-27b training into 24gb is the part worth understanding, not just trusting. is the 50% VRAM cut mostly from the FLA kernels replacing quadratic attention memory, or is gradient checkpointing doing the heavy lifting here. those two produce very different speed/memory tradeoffs at longer sequence lengths.
Nice one!
The 24GB part is impressive, but I’m more curious about the trade-off at longer context lengths. At some point VRAM savings from the kernels can get eaten up by sequence length and activations —would be really interesting to see the throughput/VRAM curve for 4K → 16K → 32K.
This is the kind of optimization that actually matters in production, more capability on practical hardware changes who can build and ship. When 27B class fine tuning fits into a 24GB workflow, the experimentation loop gets a lot more real. What are you seeing on stability and quality once these gains are pushed into real multi turn enterprise workloads?
Squeezing fine-tuning for a 27B model onto a single 24GB GPU is huge for small research teams and solo developers. When you cut VRAM overhead by 50%, you don't just save money on cloud instances—you radically accelerate the experimental feedback loop for custom domain models. Graet work !
A 27B model fine-tunable on 24GB VRAM is the part that actually changes who can experiment. Does the 50% VRAM saving hold once you push context toward the 262K limit?
Unsloth AI always contributing in the open-weights space.
In case you were wondering, Kaggle, like Colab, provides 30 hours of free GPU with 2× Tesla T4s so you can fine-tune Qwen3.8-27B completely for free by just having a Google account! 🦥 Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.