81. Minimal LLM Post-Training Experiments on an 8GB GPU在单张 8GB GPU 上运行的可复现 LLM post-training 实验,对比 SFT、DPO、GRPO …LLM大模型💬0▲0