LLaMA-Factory/examples/lora_single_gpu/README.md

6 lines
101 B
Markdown

Usage:
- `pretrain.sh`
- `sft.sh` -> `reward.sh` -> `ppo.sh`
- `sft.sh` -> `dpo.sh` -> `predict.sh`