Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
AI & ML interests
RL Environments at Scale
Recent Activity
View all activity
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 36 โข 1 -
HuggingEnvs/watercolour-reference-pool
Viewer โข Updated โข 178 โข 1.07k โข 2 -
HuggingEnvs/watercolour-rollouts-hps-only
Viewer โข Updated โข 470 โข 653 -
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 26 โข 1
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐236Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐39Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐30From an idea to a trained 4B, with the dead ends left in
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 36 โข 1 -
HuggingEnvs/watercolour-reference-pool
Viewer โข Updated โข 178 โข 1.07k โข 2 -
HuggingEnvs/watercolour-rollouts-hps-only
Viewer โข Updated โข 470 โข 653 -
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 26 โข 1
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐236Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐39Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐30From an idea to a trained 4B, with the dead ends left in