What Is OpenAI's Reinforcement Fine-Tuning?

Reinforcement fine-tuning (RFT) trains a reasoning model toward better answers on a defined task by scoring its responses. In OpenAI’s documented workflow, a custom grader supplies the reward signal. It targets measurable behavior on a specific task.
Availability update — September 8, 2026: OpenAI is winding down its self-serve fine-tuning platform. New users cannot start training, and existing users face phased restrictions. This article explains the method and its limitations; the historical demonstration below is not a current signup guide.
Is OpenAI Reinforcement Fine-Tuning Still Available?
The official deprecation schedule separates the ability to create training jobs from the ability to use an already trained model:
- May 7, 2026: organizations that had never run fine-tuning lost access to new training jobs.
- July 2, 2026: organizations that had not used a fine-tuned model for inference in the preceding 60 days also lost access to new jobs.
- January 6, 2027: remaining active customers will no longer be able to create new jobs. Existing model inference ends according to the underlying model’s retirement schedule.
Model-specific dates can arrive sooner: the same schedule lists ft-o4-mini-2025-04-16 for shutdown on October 23, 2026. Check your exact model and account eligibility; January is not a blanket guarantee of continued access. OpenAI’s Evals dashboard and API have a separate November 30, 2026 shutdown, following read-only status on October 31. Export evaluation records and plan any transition around the services you actually use.
Reinforcement Fine-Tuning in OpenAI
OpenAI’s December 6, 2024 demonstration introduces its RFT research preview, including discussion with Berkeley Lab researcher Justin Reese. It provides historical context for the training idea. Its preview models, access instructions, and program plans should not be treated as today’s offering.
RFT, Supervised Fine-Tuning, and Preference Training
These methods use different training signals. OpenAI’s model-optimization guide distinguishes them as follows. The comparison explains the techniques; availability remains subject to the retirement dates above.
| Method | Training signal | Typical purpose |
|---|---|---|
| Supervised fine-tuning (SFT) | Examples of prompts and desired responses | Teach a consistent format, classification task, or instruction-following behavior. |
| Direct preference optimization | Preferred and rejected responses to a prompt | Shape choices such as tone, style, or what a summary emphasizes. |
| Reinforcement fine-tuning (RFT) | Grades assigned to generated responses | Improve reasoning on tasks whose success can be reliably evaluated. |
Prompt engineering is another option: clearer instructions, useful context, and examples may solve the problem without a training job. Establish a baseline on representative inputs before deciding that fine-tuning is necessary. A workflow that fails because information is missing needs better information, not merely a more confident answer.
RFT Process Steps
1. Define the Task and Prepare Examples
RFT still uses a dataset. The OpenAI RFT guide describes training prompts, a grader, and separate validation data. Its documented supported reasoning-model snapshot is o4-mini-2025-04-16. Do not assume every OpenAI model supports this method.
For a hypothetical document-extraction project, define exactly what counts as correct: the requested company, reporting period, currency, and value from the supplied document. Include missing fields and conflicting figures in the task design. “Produce a convincing financial summary” is too vague to establish whether the underlying numbers are right.
2. Build and Check the Grader
A grader converts an answer into a score. OpenAI’s grader documentation describes scores from 0 to 1, including partial credit. Available designs include string checks, text similarity, model-based scoring, and Python checks; multiple graders can be combined. A separately trained reward model is not required for every RFT task.
- Exact checks: useful when a field must match a known identifier or answer.
- Code checks: useful for calculations, schema validation, or deterministic tests.
- Model-based checks: useful for rubric judgments, with expert review to detect inconsistent grading.
Consider this illustrative rubric, not an OpenAI benchmark: a supplied report states revenue of $120 million and operating income of $18 million. The requested operating margin is 18 ÷ 120 = 15%. Award 0.5 for the correct result, 0.3 for using the specified figures and period, and 0.2 for valid output structure. An answer earning only the first and third components scores 0.7. Whether that is acceptable depends on the application; it should not pass a release gate that requires all three.
3. Sample Responses, Score Them, and Update the Model
During RFT, the model generates candidate answers to training prompts. The grader scores them, and policy-gradient updates favor higher-scoring behavior. Repeating the cycle changes the model’s parameters. The guide does not specify one universal optimization algorithm.
4. Compare Training and Validation Results
Monitor training and validation rewards, including individual grader scores. The illustration below is from OpenAI’s documentation, not a LuxAlgo experiment.

Where RFT Can Help
OpenAI’s RFT use-case guide groups early applications around generating structured code that passes tests, extracting verifiable information, and applying complex rules. These are bounded objectives with assessable answers. A fluent response alone is not evidence that the task was completed correctly.
- Code generation: assess compilation, required behavior, and tests, rather than rewarding plausible-looking code.
- Information extraction: compare fields with the source and check the output structure.
- Domain rules: evaluate whether a classification or decision follows a defined policy, with appropriate expert judgment.
Medical and legal examples need specialist evaluation and a tightly defined task. They do not make an RFT model a substitute for a clinician or lawyer. Creative writing presents a different challenge: narrative quality, character voice, dialogue, and style can involve competing preferences. Decide whether a preference-based approach or better prompting fits the objective before inventing a single reward score.
ChatGPT and Codex are product names, not proof of a particular RFT training recipe. The cited documentation does not establish a universal RFT-driven percentage improvement for those products. Treat claims such as “37% more accurate” or “50% fewer unsafe outputs” as unsubstantiated unless they identify a model, dataset, baseline, metric, and reproducible source.
Benefits and Current Challenges
Targeted Improvement Depends on Good Evidence
The attraction is task-specific improvement guided by an explicit standard. However, a reported gain on one experiment does not establish broad accuracy, safety, or efficiency gains. Compare the trained model against a strong baseline on the same task, and report the costs and failure cases alongside the headline score.
OpenAI’s fine-tuning best practices emphasize data quality, consistency, diversity, and sufficient context. A small collection of carefully checked examples may be more useful than a larger noisy collection, but there is no universal promise that a dozen examples will be enough or that RFT needs 90% less data. Resolve disagreements about the desired answer before scaling the dataset.
Reward Hacking and Evaluation Gaps
A model can satisfy a weak scoring rule while missing the real goal. In our revenue example, a checker that looks only for “15%” could reward an answer about the wrong company. Check the source fields and reporting period too. Include deliberately wrong answers when testing the grader so you know what it rejects.
Keep a final test set separate from the examples used to revise the model and rubric. If every disappointing test case becomes another tuning example, the remaining score stops being a clean check of generalization. Review failures by category rather than relying on one average, and test important behaviors outside the optimized task for regressions.
Compute, Grading, and Operational Costs
Budget for more than a successful training run: dataset preparation, expert review, repeated experiments, and ongoing inference all consume resources. A slow or inconsistent grader can make experimentation expensive without delivering useful evidence. Decide in advance what improvement would justify the work and when to stop. Current platform retirement dates also make portability and record retention part of that decision.
Applying the Evaluation Mindset to AI Trading Research
For a trader, the useful connection is disciplined testing of AI output. LuxAlgo is a charting and AI platform where Quant, our coding agent, can turn defined trading rules into a strategy script. Open Code to inspect it, then sign in and Run it on the intended chart. This is an AI-assisted strategy workflow, not a claim that LuxAlgo trains its models with OpenAI RFT.
For example, ask Quant to test a specific moving-average crossover with explicit exits and position sizing. Before judging returns, verify that entries and exits match the written rules. A script that runs successfully has passed a technical check; whether the trading idea is useful remains a separate research question.
The strategy viewer supports reviewing net profit, trade count, drawdown, profit factor, and the Trades Log. Configure commission and slippage, inspect individual trades, and compare results across an untouched period. An attractive backtest is not a reward function that proves future trading success.
The same question matters in both settings: does the measured score represent what you actually need? Preserve the rules, data assumptions, and evaluation records so you can explain a result and recognize when it no longer applies. Use AI to make the research process more inspectable, while keeping the final decision tied to evidence.
References
- OpenAI: deprecations and fine-tuning availability
- OpenAI: reinforcement fine-tuning guide
- OpenAI: graders
- OpenAI: model optimization and fine-tuning methods
- OpenAI: reinforcement fine-tuning use cases
- OpenAI: fine-tuning best practices
- OpenAI: December 2024 RFT demonstration
- LuxAlgo: making strategies with Quant
- LuxAlgo: strategy viewer and backtest settings
Read next