Training your own rat model
Overview
From SDK version 0.4.0, a researcher's own model can control the rat instead of RODENT's built-in statistical rat. The model practises in a RODENT experiment again and again, learns from a reward the researcher defines, and is then saved to the project so the whole team can run it in their experiments and replay its runs on the website.
Training happens in Python, usually Google Colab, with PyTorch. From SDK version 0.4.3, saving a compatible network also stores an ONNX copy. The website can run that copy live from Setup > Rodent > Who controls the rat, or run a researcher-uploaded .onnx model. It can still replay runs produced in Python under 3 Review.
Two things stay true whatever model is used:
- The model's rat obeys exactly the same walls, doors, arena edge and speed limit as the built-in rat. Both use the same movement code (
RatBody). - A trained model is a piece of software that learned to score well on a reward. It is not evidence of how a real rat behaves. Report the reward, the arena and the training settings with any result.
What you need
- A RODENT website account and a project. To save a model you must be the project lead or an editor; any member can load and run saved models.
- A Google account for Colab, or any Python 3.9+ with
pip install "rodent-sdk[train]"and PyTorch. - The training notebook. On the website, open the experiment, choose Use in Python, and click Open the training notebook. Copy the experiment code from the same dialog.
Five words used throughout
| Word | Meaning |
|---|---|
| Episode | One attempt. The rat is placed in the arena and gets a fixed number of steps (10 steps = 1 second of rat time). Training is hundreds or thousands of episodes. |
| Observation | What the rat senses before each step: 31 numbers between -1 and 1 (see below). |
| Action | The move the model makes on each step. |
| Reward | A number after each step that scores it against the goal. Training changes the model so it collects more reward. |
| Round | A batch of episodes played with the current model, after which the model learns from them. |
What the rat senses (the observation)
| Group | Values | Meaning |
|---|---|---|
| Body | x, y, heading_sin, heading_cos, speed, last_turn | Where the rat is, which way it faces, how fast it moved and how much it last turned. |
| Walls | wall_ahead, wall_ahead_left, wall_left, wall_behind_left, wall_behind, wall_behind_right, wall_right, wall_ahead_right, touching_wall | Distance to the nearest wall or closed door in 8 directions (0 touching, 1 at 50 cm or more) and whether it is touching one. |
| Light, sound, odour | for each: strength, strength on the left side, on the right side, and the direction it comes from (sine and cosine relative to the heading) | Light travels in straight lines and stops at walls and closed doors; sound and odour weaken through walls; odour spreads over time as a gas. |
| Time | episode_progress | How far through the episode the rat is (0 to 1). |
The names, in order, are in env.observation_names.
What the rat can do (the action)
Two action styles are available. The notebook uses discrete:
| Style | Actions |
|---|---|
actions="discrete" | 0 stay, 1 forward, 2 turn left, 3 turn right, 4 run |
actions="continuous" | two numbers: turn from -1 to 1 (up to 45 degrees a step) and speed from 0 to 1 (up to the built-in rat's fastest step) |
The training notebook, step by step
The notebook is written as a worked example: aim, method, training, results, saving and reuse. Cells marked ✏️ are the ones a researcher edits; everything else runs as it is.
| Step | What happens |
|---|---|
| 1. Setup | Installs rodent-sdk[train] and downloads the example model file rat_model.py next to the notebook. |
| 2. Sign in | Uses the RODENT_EMAIL and RODENT_PASSWORD Colab Secrets, or asks. |
| 3. Choose the experiment ✏️ | Paste the experiment code from Use in Python, or type the project and experiment names, or create a new Light/Dark Box project. An optional cell changes the setup (lights, odour, sound, doors) and saves it. |
| 4. Explore the experiment | Prints the regions, doors and stimuli (on, off or not enabled), draws heat maps of what the rat can sense, and runs the built-in rat as a baseline. |
| 5. Define the task ✏️ | Choose the goal (TASK), the start position, the episode length and how long to train. |
| 6. The model | Loads RatModel from rat_model.py and prints its settings. |
| 7. Train | Tests the untrained model, then trains round by round and prints progress. Each run of the training environment cell starts its own log folder. |
| 8. Results | Learning curves, the before and after test, any episode rebuilt and animated, the first, middle and last episodes side by side, and the trained rat next to the built-in rat. |
| 9. Save to the website ✏️ | Saves the model as a new version, and the trained rat's run and a chosen episode to Review, with links. |
| 10. Reuse ✏️ | Runs the saved model in another experiment of the same project. |
| 11. Download | The model file, its weights, the training history and every logged episode. |
Choosing a task
TASK | What the rat should learn | Reward after each step | Measure |
|---|---|---|---|
reach_region | get to the region GOAL | +1 inside the goal; before that, a small minus that shrinks as it gets closer | % of episodes ending in the goal |
avoid_light | stay out of the light | minus how bright it is | % of time spent in the dark |
follow_odour | find the odour source | how strong the odour is | odour strength over the last fifth of each episode |
follow_sound | find the sound source | how loud the sound is | loudness over the last fifth of each episode |
A researcher can also replace the reward with any function of step.region, step.position, step.hit_wall, step.moved, step.light, step.sound and step.odour.
Results from our own runs of the notebook, on the same 20 test episodes before and after training (80 rounds of 8 episodes):
| Task and arena | Before training | After training |
|---|---|---|
reach_region, Light/Dark Box, goal dark_room, start next to the lamp | 0% of episodes | 100% |
avoid_light, Light/Dark Box | 14% of time in the dark | 87% |
follow_odour, Light/Dark Box with the odour source switched on, 300-step episodes | odour strength 0.03 | 0.77 |
follow_sound, Complex Habitat with a 50 cm sound radius | loudness 0.02 | 0.10 (learns slowly) |
For avoid_light, the trained rat waited in an unlit corner of the bright room, outside the lamp's reach. That scores as well as reaching the dark room, which is a reminder to read every measure against the reward that produced it.
Sources must be enabled and active
A light, odour or sound source works only if it is enabled (present in the arena) and active (switched on at the start). The templates ship their odour and sound sources disabled. Switch one on in the website's arena editor (tick Enabled and Active at start, then save), or in the notebook's optional setup cell:
exp.stimuli["dark_odor"].enabled = True
exp.stimuli["dark_odor"].active = True
exp.stimuli["dark_odor"].radius = 30 # how far it reaches, in cm
exp.save()
Training always uses the experiment as it is saved, so the notebook saves a changed setup before it trains. Sound is heard only inside its radius (18 cm in the templates), so a larger radius helps a rat that starts far away. Odour needs time to spread, so longer episodes help.
The model file
The example model lives in its own file, rat_model.py, so it can be read and replaced without touching the notebook. It is a PPO actor-critic (Schulman et al., 2017):
- the actor, a small neural network, scores the 5 moves from the 31 observations, and the rat picks moves in proportion to those scores, so it keeps exploring while it learns;
- the critic, a second network, estimates how much reward is still to come from where the rat is;
- after each round, every move is judged by whether things went better than the critic expected (the advantage); better moves become more likely, and PPO limits how far the actor can change in one round so training stays stable.
Bringing your own model
Replace rat_model.py with a file that defines RatModel, a PyTorch module with three methods:
class RatModel(torch.nn.Module):
def forward(self, observations):
"""31 numbers in, a score for each of the 5 moves out."""
def act(self, observation, best=False):
"""The move to make, 0 to 4. RODENT also calls this when the model drives a run."""
def learn(self, episodes):
"""Improve from finished episodes. Each has observations, actions, rewards,
final_observation and terminated. Return a dict of numbers to print."""
When a model with an act method drives a normal run, RODENT asks it for each move exactly as in training, and seeds PyTorch from the experiment's seed, so the same run comes out every time. A network without act is driven by its highest-scoring move.
Looking at any episode
Every training episode is logged with its seed and moves, so any episode can be rebuilt exactly afterwards, without the model:
episode = env.log.episode(522) # episode 522, with its full recording
episode.plot(steps=(0, 150)) # its path
episode.at(50) # what the rat sensed and did at step 50
episode.animate() # an animation in the notebook
env.log.compare([1, 240, 480]) # paths side by side
env.log.plot() # learning curve and time per region
Episodes are numbered from 1, and episode n uses the experiment's seed plus n. Moves are stored at 16-bit precision and sent to the simulator already rounded, so a rebuilt episode matches the trained one step for step.
Saving and sharing
version = project.save_model("Dark seeker", model, notes="...", training=env.log)
model_run.save() # the trained rat's run, under Review
episode.save() # a training episode, under Review
print(episode.review_link(50)) # opens the replay at step 50
- A model belongs to a project. Saving again with the same name adds version 2, 3 and so on. Versions never change after upload, so a saved run always points to the exact weights that produced it.
- Each version records how it was trained: the experiment, the number of episodes, the learning curve and the simulator version.
- Weights files are stored privately (up to 50 MB each) and are always loaded with PyTorch's safe
weights_only=Truemode. - An episode can only be saved to Review while the experiment on the website is still the setup it was trained on; otherwise the replay would show the rat in the wrong arena.
What appears on the website
- On the project page, Trained rat models lists each model with its latest version, how many episodes it trained for and in which experiment, its notes, a small learning curve, the average reward of its last 100 training episodes, and a Copy Python to load it button.
- Under 3 Review, a saved run driven by a model is labelled with the model and version (for example "Dark seeker v1"), and a training episode with its number. A review link opens the replay at the chosen step.
Reusing a model
Every experiment in the project can use every model in the project:
model = project.model("Dark seeker").latest.load_into(RatModel()) # or .version(2)
run = project.experiment("Door closed").run(steps=300, model=model, episode_length=150)
run.save() # under Review for "Door closed", labelled "Dark seeker v1"
Typical uses are comparing one trained rat across several experiments (a door closed, a brighter light, a treatment), or loading a model and training it further, then saving the result as the next version.
The model reacts to what it senses and to where it is, so it transfers well to experiments with the same arena layout and less well to a very different arena, where it is a starting point for more training. To use a model in another project, download a version (version.download("weights.pt")) and save it there (other_project.save_model("Dark seeker", "weights.pt")); the copy starts at version 1 without the original training record.
Without the notebook
The training environment follows the Gymnasium format, so any Gymnasium-style training loop works:
import rodent_sdk as rodent
exp = rodent.experiment("light_dark_box", seed=1) # or a website experiment
env = exp.training_env(max_steps=150, reward=my_reward, actions="discrete",
start={"random": True, "position": [20, 50]}, log="my-log")
obs, info = env.reset()
obs, reward, terminated, truncated, info = env.step(1) # forward
copies=n runs several simulators side by side. One simulator runs about 1,000 steps a second; three side by side ran about 1,500 a second on a two-core machine. The full list of calls is in Python SDK and Headless API.
Rules and limits
| Situation | What happens |
|---|---|
| Training on an experiment with unsaved changes | Refused. Save the experiment first. |
| A log folder made with a different experiment or settings | Refused. Each training run uses its own folder (the notebook makes a new one each time). |
| Saving a model without lead or editor access | Refused by the database. |
| A weights file over 50 MB | Refused by storage. |
| A training episode after the experiment changed | Can no longer be saved to Review. |
| A saved model with a different network shape | load_into fails; the network must match the one that was saved. |
Something went wrong?
- The measure does not improve: train longer (run the training cell again), check in step 4 that the rat can sense what the task needs from where it starts, and make sure the reward says "warmer, colder" on the way to the goal.
- "needs an odour/sound source that is enabled and active": see Sources must be enabled and active.
- "The log ... was made with a different experiment": run the training environment cell again; it starts a new log folder.
- "The simulator would reject this setup": for example two sources with the same id. Nothing was saved; run step 3 again and choose a new id.