Tech article

Learning to solve hard problems in RL for LLMs by never giving up

No preview is available. Read the original article for the full story.

Hacker News | Sep 15, 2026 | natolambert

Automated excerpt

The hardest problems are barely improving. NGU allows $k=4$ to keep retrying hard problems, solving them as well as $k=32$ early in training. RL improves easy tests to nearly fully solved.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More tech news