alias

OpenAI::FineTuneDPOHyperparametersBeta

The beta value for the DPO method. A higher beta value will increase the weight of the penalty between the policy and reference model.