Пока нет отзывов
5 000 ₽
Reinforcement Learning for LLM
Открыть наSTEPIK.ORG
An engineering course on reinforcement learning for LLMs and agentic systems: from MDPs, bandits, REINFORCE/PPO, and offline RL to reward models, RLHF/DPO, RLVR/GRPO, test-time search, agentic RL, infrastructure, and evaluation. In addition to the open lectures, the paid Stepik course includes more than 500 tasks: auto-graded tests, exercises with solutions, self-check questions with answers, and notebooks. The course is part of an evolving series on ML and LLMs.
| Показатель | Текущие показатели | Рост | |||
|---|---|---|---|---|---|
| Значение | 🏆 Рейтинг | 3 дн | 7 дн | 30 дн | |
| 0 | |||||
| 0 | |||||
| 0 | |||||
| 0.000 | |||||
| 114 | |||||
| 136 | |||||
| 7 | |||||
| 5 000 ₽ | — | ||||
| — | — | ||||
| — | — | — | — | ||
| — | — | — | — | ||
| Сложность | normal | — | — | — | — |