Background.

(A) Recent work has identified explicit sequence representations in prefrontal cortex during working memory (Xie et al., 2022; El-Gaby et al., 2023). When an animal has to execute a behavioural sequence (left), individual neurons represent conjunctions of location and sequence element (centre). Separate populations (planes) therefore represent the expected location at different times in the future. The entire sequence is represented concurrently by the simultaneous firing of different neurons that encode the expected location at each time in the future (right). (B) An example V1 cell fires when its receptive field is aligned with an inferred line (blue), but not for a control stimulus with no inference (orange). Such visual inference is mediated by structural priors embedded in the circuit connectivity, where neurons representing consistent visual features excite each other (Iacaruso et al., 2017; Shin et al., 2023). Figure adapted from Lee and Nguyen (2001). (C) Structural knowledge is embedded in the synapses of the head direction system (centre; Turner-Evans et al., 2020), which constrains the network to represent single angles (Kim et al., 2017). Visual and proprioceptive inputs (left) determine which angle should be represented (right). We suggest that a mechanism like (B) and (C) infers the representation in (A). (D) Prefrontal cortex is particularly important in dynamic environments. Stimulus-response associations and repeated choices are robust to prefrontal lesions (i; ii), and acortical mice can solve spatial memory tasks (iii; Zheng et al., 2024). However, PFC is needed for reversal learning (iv; Walton et al., 2010) and when different goals are important, or ‘rewarding’, at different times (v; Shallice and Burgess, 1991). In multiplayer board games (vi), different resources are valuable at different stages, and opponents can dynamically change their strategy.

The spacetime attractor.

(A) Left; in a one-dimensional ring attractor, neurons representing similar directions excite each other (top, red), and different directions inhibit each other (blue). Connections are shown for the green example cell. Stable network states (fixed points) have neurons at a particular encoded angle being most active (bottom). Right; grid cells (bottom) emerge from a two-dimensional attractor network (top), where cells with similar preferred locations excite each other (red) and intermediate distances inhibit each other (blue). (B) The spacetime attractor encodes a three-dimensional representation, where neurons have both a preferred location (planes) and delay (horizontal axis). This resembles the PFC representation in Figure 1A. Cells that represent adjacent locations in both space and time excite each other (top, red), and different locations at the same time inhibit each other (blue). Given inputs indicating the current location (mouse) and future reward (cheese), the fixed points represent reward-maximising paths through space and time (bottom). While this example has a single static goal, reward inputs can also differ between subspaces if the reward is dynamic. (C) Initial activity is diffuse (left), followed by convergence to a stable representation of a plan (centre). At convergence, a policy is read out from the subspace that represents the next location along the desired trajectory. When the agent moves to the next state, neural activity updates to represent the remaining plan-to-go (right; El-Gaby et al., 2023).

STA dynamics and fixed points.

(A) STA dynamics on a 3 × 3 grid with a single goal (top). The final plane in each row is a max projection that summarises the encoded path. Time is measured in units of the network time constant τ. The converged representation is a stable fixed point that resembles a posterior distribution over trajectories conditioned on reaching the goal (Methods). (B) In an environment with several disjoint paths to the goal, there is initial competition between representations of the alternatives. The factorised representation then drives convergence to a single trajectory (Methods). (C) The STA converges to different fixed points on different trials in the environment from (B). (D) There is not a strong bias towards either of the two fixed points. (E) If the reward input is set to 0 after convergence in (B), the dynamics relax back to a ‘diffusive’ representation around the initial location. This illustrates the input-dependence of the fixed points. (F) In an environment with both a shorter and a longer path to the goal (top), the STA reliably converges to a representation of the shorter path (centre). This is because the shorter path has a higher likelihood and larger basin of attraction in the network dynamics (Methods). If the network is initialised close to a representation of the longer path, it converges to this solution instead (bottom).

Model comparisons in static and dynamic tasks.

(A) Example tasks with different degrees of dynamic reward. Colours indicate reward at each location (white to green), and red arrows indicate the optimal policy. In the ‘moving goal’ task, the agent intercepts a target that moves along a different trajectory in each trial (arrow). In the ‘reward landscape’ task, the reward is sampled independently in space and time. The agent has to maximise cumulative reward, and it can be optimal to forgo immediate reward for a later payoff. (B) Representations and performance in the static goal task with a fixed goal. Left: value functions learned by TD and SR agents. Centre: STA representation encoding a path through space and time. Right: performance of each model (see Figure S1 for STA performance in larger environments). The grey bar is a ‘greedy’ baseline that maximises immediate reward, and the horizontal line is a random baseline. (C) Task with a static goal that changes between trials. (D) Moving goal task. Left: the SR computes a value function that averages reward across time-within-trial. Middle: the STA takes into account the moving goal and computes a path that intercepts it. Right: performance of each agent. (E) Performance in the reward landscape task. All error bars indicate 1 standard deviation across 20 agents and environments (dots).

RNNs learn spacetime representations in the reward landscape task.

(A) An STA infers the entire future during planning and represents the ‘future-to-go’ during execution (left). The representation in each subspace generalises across all trajectories passing through that location in spacetime (right). (B) We compare a spacetime attractor; an RNN trained on the reward landscape task; and an agent that computes an exact value function in space and time. The value-based agent computes an optimal policy from ‘neural activity’ containing (i) the value function, (ii) the agent location, and (iii) the time-within-trial (Methods). (C) Decoding accuracy at the end of planning for: left; agent location at each time in the future. Right; the time at which the agent will be at a given location, plotted as a function of the actual time the location was visited. Decoders were trained in crossvalidation across the current agent location (Methods). (D) Decoding accuracy during execution of location at each time in the past or future. (E) We trained a single decoder to predict location at time 3 from neural activity at time 1 (black circle). The same decoder predicted location at time t + 2 from neural activity at any other t, demonstrating ‘conveyor belt’ dynamics. (F) Performance in the static goal (left), moving goal (centre), and reward landscape (right) tasks for RNNs trained on either task (x-labels; colours). (G) Normalised parameter magnitudes of the three RNNs (left) and average firing rates in the static goal task (right). All error bars indicate 1 standard deviation across 5 RNNs (dots).

RNNs learn a world model.

(A) The weights of RNNs trained on the reward landscape task are projected into an orthonormal coordinate system with axes that predict different points in spacetime (top). The projected weights are interactions between points in spacetime (bottom). (B) The average recurrent weights between subspaces separated by a single action resemble the environment adjacency matrix (bottom, ‘RNN’). ‘Empirical STA’ is the same analysis performed on approximate subspaces estimated from neural activity in the handcrafted STA. ‘True STA’ indicates weights between the ground truth future-coding subspaces in the handcrafted STA. Green box in ‘True STA’ indicates weights between (i) the single location denoted by a green circle in ‘Prediction’, and (ii) all locations in the subspace indicated by a green square. (C) Input weights to the ‘current’ (δ = 0; left) and a ‘future’ (δ = 2; right) subspace. Consistent with an STA, the current subspace of the RNN receives location input, and future subspaces receive reward information. (D) Average recurrent weights between an example location (green arrow) and other locations at increasing time differences (Δ; planes). Locations in subspaces Δ apart are more strongly connected if they can be reached in Δ actions. Light green circles indicate these Δth order adjacency matrices (Δ = 0 corresponds to the identity). (E) Correlation between (i) the average connectivity between subspaces separated by Δ actions (lines; legend), and (ii) different order adjacency matrices (x-axis). Shading indicates standard deviation across 5 RNNs. (F) Example environment with a high-value path and a lower-value path. (G) Weak stimulation does not affect the RNN spacetime representation, but strong stimulation switches it to the lower-value path. (H) Magnitude of RNN representational change across time and stimulation strength (Methods).

Spacetime attractors can adapt to changing structure.

(A) Changing structure requires adaptation. (B) Performance of RNNs trained on the reward landscape task in a single maze or with a different structure on every trial, when evaluated in either a single maze (top) or across changing mazes (bottom). (C) Two example mazes (left) and their corresponding adjacency matrices (centre). The effective connectivity between adjacent subspaces in a single RNN (right) reflects the structure of each maze. (D) Average correlation between subspace connectivity in a maze and the structure of either the same (blue) or a random (grey) maze. (E) The spacetime representation is more similar across distinct trials from the same maze than trials from different mazes (left). By using slightly different subspaces in different environments, a network can match its connectivity to the environment structure (right; orange vs. blue). (F) Putative mechanism for structural generalisation. Instead of representing each future location, neurons encode expected future transitions (black arrows in example state). Each transition (green) excites other transitions that can follow or precede it (red arrows). Structural input to PFC inhibits transitions that are not available in a given environment (blue), preventing planning between states that would otherwise be connected (light red arrows). (G) Effective connectivity between directions in neural state space that encode ‘consistent’ consecutive future transitions ( and ; green to red), ‘adjacent’ transitions (green to blue), or any other transitions (green to grey). (H) Projection of the input from wall wij onto representations of future transitions through the wall (; dark blue), transitions to other adjacent states (; light blue), or any other transitions (grey). All error bars indicate 1 standard deviation across 5 RNNs (dots).

STA performance for different planning horizons.

(A) We considered different 6 × 6 mazes with a single static goal location (cheese). We then quantified how reliably the STA inferred the shortest path to the goal as a function of the initial distance from the goal (mice). (B) Fraction of problems where the STA infers the shortest path, as a function of the distance from the goal. This analysis was repeated for four different noise levels: (i) No noise (blue). (ii) The baseline noise used in the main text (Methods; yellow). (iii) Twice the baseline noise. (iv) Three times the baseline noise. All simulations used an input strength of β = 6 (Methods), and representations were analysed after 1500 network iterations (30 time constants) to ensure convergence. Shading indicates standard error across 20 different environments. When the STA fails to infer the shortest path, the predominant failure mode corresponds to representations of ‘discontinuous’ paths that move immediately from the start location to the goal. This can happen if the reward inputs dominate over the recurrent signal, which decays with planning depth for intermediate locations. It can also happen if noisy connectivity leads to spurious connections between distant locations in adjacent subspaces. (C) The failure mode of discontinuous representations can be mitigated if the distance to the goal is known a priori. In this simulation, reward input was only provided beyond the first time at which the goal is reachable in a 10 × 10 maze with no noise. A higher input strength of β = 20 can then be used, and simulations were run for 300 time constants to ensure convergence. The STA reliably plans trajectories consisting of 15 actions, and it sometimes infers shortest paths that are 30 actions long. (D) We do not think planning depth is a major limitation of the STA as a model of human and animal behaviour, because we rarely plan more than 6-7 steps ahead at a single level of abstraction (van Opheusden et al., 2023). Instead, long-horizon planning can be solved hierarchically (Eckstein and Collins, 2020). In this example, an agent has to plan a trajectory of 40 actions (red arrow) in a large 16 × 16 environment. However, the environment has been grouped into 16 separate 4 × 4 regions (colours). (E) We implemented a simple hierarchical spacetime attractor. The first ‘abstract’ STA knows the transition structure between regions, and it infers a sequence of seven regions to the goal. The second ‘detailed’ STA receives the boundary states of the inferred next region as a reward input (red arrows). It infers a sequence of seven primitive states to the next region. Once the second region is reached, the abstract STA updates its representation. This allows the detailed STA to plan its way to the third region. Hierarchy enables planning of trajectories that are exponentially long in the depth of the hierarchy, and therefore in the number of neurons.

Recurrent weights optimised for planning-as-inference.

(A) Planning can be implemented as an inference process over future locations. The posterior at each point in time is given by the product of a ‘forward message’, a ‘backward message’, and a ‘likelihood’ that reflects the reward at different points in space and time. (B) We trained a network with this product-of-messages architecture to learn fixed points that approximate the true posterior (Derivation). The fixed points are given by where {G} are learnable parameters and eβR is the likelihood. (C) Average KL divergence between the true and approximate marginals over the course of learning. (D) Learned connectivity between all pairs of subspaces. The lower triangular blocks correspond to , and the upper triangular blocks to . (E) Left: 1st, 2nd, and 3rd powers of the transition matrix. Centre: average forward connectivity between populations separated by 1, 2, or 3 actions. Right: average backward connectivity between populations separated by 1, 2, or 3 actions. (F) Correlation between (i) the average connectivity between subspaces separated by Δ actions (lines; legend), and (ii) different powers of the transition matrix (x-axis). (G) Average normalised connectivity strength as a function of the number of actions separating two populations, evaluated separately for (blue) and (orange). This analysis excluded connections to and from the δ = 0 population, which has a stronger likelihood term than all other populations. Lines and shading in (F) and (G) indicate mean and standard error across 5 networks trained in different environments.

Schematic illustration of model inputs and outputs.

(A) The RNNs and STA received three different kinds of inputs: (i) the current agent location; (ii) the scalar reward available at every location at every time in the future; and (iii) a binary vector indicating the presence or absence of barriers between otherwise adjacent states. The wall input was provided to all models, but it is only relevant for the RNNs trained in Figure 7, where the environment structure changes across trials. The output was an ‘allocentric’ policy, indicating the desirability of moving to each location in the environment. This policy was normalised over locations adjacent to the current location before sampling an action. (B) For all analyses of the handcrafted STA and the RNN in Figure S10, reward inputs were provided continually and in ‘relative time’. In other words, a consistent input channel indicated what the reward would be at a particular location in δ actions. (C) All other RNNs were trained in a ‘working memory’ setting, where reward inputs were only provided during the initial ‘planning’ phase. During subsequent execution, all reward inputs were set to 0.

RNN performance during and after training.

Each line in this figure corresponds to one of the five RNNs that were used for analyses in the main text. (A) Performance over the course of training, averaged over all actions within each trial. (B) Performance at the end of training as a function of the action number within the trial. When assessing performance at time t, trials were only included that had optimal choices up to time t − 1. (C) Probability of choosing the action with highest value as a function of the value difference between the two actions with highest value. This analysis shows that errors are only made then the optimal action is close in value to the second best action. (D) Probability of choosing the action with highest reward as a function of the reward difference between the two actions with highest reward. Reward is less predictive of behaviour than value, confirming that the RNNs compute long-term value rather than relying on greedy reward.

Additional analyses of learned RNN representations.

(A) We compare a spacetime attractor; an RNN trained on the reward landscape task; and an agent that analytically computes a full spacetime value function. The value-based agent computes an optimal policy from ‘neural activity’ containing (i) the value function, (ii) the current location, and (iii) the time-within-trial (Methods). (B) Decoding accuracy of agent location at different times (x-axis) from neural activity at every other time (lines; legend). All decoders were trained in crossvalidation across the current agent location (Methods). This is why the accuracy is zero when decoding location from activity at the same time. (C) We trained a single decoder to predict location at time 3 from neural activity at time 1 (green circle). The same decoder predicts location at time t + 2 (x-axis) from neural activity at any other t (lines). (D) Similarity of decoding patterns to idealised representations of future location in ‘relative’ or ‘absolute’ time (schematics).

Characterisation of subspaces learned by the reward landscape RNN.

We analyse the orthogonal subspaces representing future behaviour that were used to characterise RNN connectivity and dynamics in Figure 6. (A) Left: Empirical overlap between future-coding subspaces estimated during the planning and execution periods of the RNN. The RNN uses separate subspaces for computation of a plan and subsequent execution. Right: Variance explained by the five 16-dimensional subspaces that predict the current and next four agent locations, estimated during planning (orange) or execution (blue). The variance explained was quantified separately at each time-within-trial, which shows information transfer from the planning to the execution subspaces at movement onset (vertical line). The grey line indicates the variance explained by the top 80 PCs, which were computed separately for each time-within-trial and therefore represent a strict upper bound. (B) As in (A), now for a network trained with continual reward input throughout the trial instead of only during planning (Figure S3; Figure S10). (C) Left: cumulative normalised variance explained by the first 80 principal components of the handcrafted STA (x-axis) at different times during the execution period (lighter to darker lines). The representation becomes lower dimensional throughout execution, because the plan-to-go becomes progressively shorter. This was also true in the task-optimised RNNs (centre & right). The RNN representations were generally high-dimensional, and they had participation ratios of 50-100 during late planning and early execution. (D) Left: normalised variance explained by each empirical 16-dimensional execution subspace in the handcrafted STA (x-axis) at different times during the execution period (lighter to darker lines). All subspaces are active during early execution, while only the ‘immediate future’ subspaces are active during late execution. A similar trend was seen in the task-optimised RNNs (centre & right).

RNNs trained on simpler tasks do not learn spacetime representations.

In this figure, we analyse the representations of four RNNs trained on all combinations of (i) the static goal or the moving goal task, and (ii) reward input throughout the task (‘continual’) or reward input only during the planning phase (‘working memory’). For all analyses in this figure, we only included trials where an optimal agent would intercept the goal in 3 to 6 actions. (A) Decoding accuracy for agent location at different times (x-axis) from neural activity at the end of the planning period. Decoders were trained in crossvalidation across the current agent location. Only the RNN trained on the moving goal task in a working memory setting learns a generalisable representation of the future. This network is also unlikely to have learned a full STA, since it fails catastrophically on the reward landscape task (Figure 5F). Note that the decoding accuracy generally increases slightly for the true STA as a function of time-within-trial. This is because the navigation tasks have strong correlations between consecutive positions, which leads to overfitting on the training data. This overfitting is less prominent later in a trial, where the space of possible locations conditioned on the current location is larger. (B) In this analysis, we trained a decoder to predict whether the agent would be at a particular location at any time in the trial from neural activity at the end of planning. Binary decoders were trained for each possible future location in crossvalidation across the current agent location. Bars indicate the average predictive accuracy across all binary decoders and current locations. The simpler networks learn representations of whether they will be at a given location at some point in the future.

Parameters learned by the reward landscape RNN.

Network weights are projected into an orthonormal coordinate system with axes that maximally predict future locations. All weight matrices are for a single example RNN, since the environment differs between networks, and the connectivity is therefore slightly different. (A) Structure of the environment that the RNN was trained in, illustrated as the 0th order adjacency matrix (the identity matrix), the 1st order adjacency matrix, and the 2nd order adjacency matrix (B) Recurrent weight matrix estimated during the planning period, which resembles the environment adjacency matrix in the off-diagonal blocks. (C) Recurrent weight matrix estimated during the execution period, which has an additional ‘feedforward’ component that transfers information from later to earlier subspaces. We expect this component of the connectivity matrix to help implement the conveyor belt dynamics identified in Figure 5E (Supplementary Note). (D) Input weight matrix estimated during the planning period. The ‘current’ subspace receives location input, and future subspaces receive reward information corresponding to the appropriate time. (E) Output weight matrix estimated during the execution period. The policy is read out from the ‘immediate future’ subspace.

Additional analyses of attractor dynamics.

(A) Change in implied spacetime representation over time in the trained RNN for different perturbation strengths (reproduced from Figure 6H). (B) Change in spacetime representation over time in the handcrafted spacetime attractor for different perturbation strengths. In contrast to the trained RNN, the ‘low value’ path is a fixed point of the perturbation-free network dynamics in the handcrafted network. At the end of a sufficiently strong perturbation, the representation can therefore stay in this new fixed point. (C) Change in RNN firing rates for different perturbation strengths. While the change in implied spacetime representation completely saturates with perturbation strength, the change in firing rates continues to increase with perturbation strength. This is expected because the network has a non-saturating ReLU nonlinearity. Small external perturbations are still quenched when quantifying the change in representation using the raw firing rates instead of the implied spacetime representation.

RNNs trained with continual reward input also learn spacetime attractors.

In the main text, we focused on an RNN that was trained with reward input provided during an initial ‘planning phase’, while no information was given about the reward function during subsequent ‘execution’. In this figure, we perform some of the same analyses on an RNN that receives reward input throughout the entire task. In this setting, the task could in theory be solved using a ‘feedforward’ strategy that does not rely on recurrent dynamics at all. However, the RNNs still learn a spacetime attractor-like solution. (A) Decoding accuracy of agent location at different times (x-axis) from neural activity at every other time. Each line corresponds to predictions from neural activity at a different time in the trial from t = 0 (yellow) to t = 4 (blue). Decoders were trained in crossvalidation across the current agent location. (B) We trained a single decoder to predict location at time 3 from neural activity at time 1 (green circle). The same decoder predicts location at time t + 2 (x-axis) from neural activity at any other t (lines). (C) Similarity of decoding patterns to idealised representations of future location in ‘relative’ or ‘absolute’ time (Figure S5). (D) The average recurrent weights between subspaces separated by a single action resemble the adjacency matrix of the environment. (E) Correlation between (i) the average connectivity between subspaces separated by Δ actions (lines; legend), and (ii) different order adjacency matrices (x-axis). (F) Input weights to the ‘current’ subspace (δ = 0). (G) Input weights to a ‘future’ subspace (δ = 2).

RNNs with a local action space also learn spacetime attractors.

In the main text, we analysed an RNN that generated a global allocentric policy, consisting of a probability distribution over all locations that was renormalised over ‘adjacent’ locations before sampling an action. In this figure, we perform some of the same analyses on a network that outputs a ‘local’ policy in an action space consisting of ‘north’, ‘south’, ‘east’, ‘west’, and ‘stay’. This RNN also learns a spacetime attractor. (A) Decoding accuracy of agent location at different times (x-axis) from neural activity at every other time. Each line corresponds to predictions from neural activity at a different time in the trial from t = 0 (yellow) to t = 4 (blue). Decoders were trained in crossvalidation across the current agent location. (B) We trained a single decoder to predict location at time 3 from neural activity at time 1 (green circle). The same decoder predicts location at time t + 2 (x-axis) from neural activity at any other t (lines). (C) Similarity of decoding patterns to idealised representations of future location in ‘relative’ or ‘absolute’ time (Figure S5). (D) The average recurrent weights between subspaces separated by a single action resemble the adjacency matrix of the environment. (E) Correlation between (i) the average connectivity between subspaces separated by Δ actions (lines; legend), and (ii) different order adjacency matrices (x-axis). (F) Input weights to the ‘current’ subspace (δ = 0). (G) Input weights to a ‘future’ subspace (δ = 2).

RNN representations and performance across network sizes.

We trained a series of RNNs with different network sizes ranging from 50 (dark blue) to 800 (yellow) hidden units. (A) Networks with approximately 300 or more units learned a spacetime representation, and the future could be decoded from the hidden state of the network at the end of the planning period. (B) Task performance saturated as a function of network size at approximately 300 hidden units. (C) Task performance increased with the ability of the network to represent the entire future explicitly. These results mirror the findings of Whittington et al. (2023) that RNNs trained on working memory tasks learn a similar ‘slot-like’ solution only if the network is large enough.

Additional analyses of RNNs trained with changing environment structure.

(A) Learning curves of RNNs trained on the reward landscape task in a single maze (‘fixed maze’; orange) or with a different structure on every trial (‘changing maze’; blue). (B) We took the RNNs trained across many mazes and evaluated them in a single maze (blue). Future states could be decoded almost as well as in the RNNs trained in a single maze (orange). Additionally, the representations in each maze were more similar to idealised representations of future location in ‘relative’ than ‘absolute’ time (right). (C) When considering data across many mazes, future transitions could be decoded in a way that generalised across current transition and maze structure (blue). Such a representation did not exist in RNNs trained in a single maze (orange).