Machine learning · A practical field guide
Policies, Values, and Planning: Understanding Reinforcement Learning
What the methods learn, how they choose, and why you might prefer one over another.
A robot reaches a junction in a maze. Left looks familiar. Right might be a shortcut. Which way should it go?
One system compares learned scores for left and right. Another has learned a rule that directly picks a direction. A third imagines several possible futures before moving. All three can be part of reinforcement learning.
The confusing part is that the names describe different things. “Actor–critic” describes the roles of learned components. “Off-policy” describes which experience an update can use. “PPO” names a particular algorithm. They belong on different lines of the same description.
We will keep returning to this robot. Each new method should answer two concrete questions: what does it learn, and what determines its next action?
1. Learning from consequences
In supervised learning, a training example often comes with the answer: this image contains a cat. In reinforcement learning (RL), the learner takes an action and receives feedback about what happened. It usually gets no label saying which action would have been best.
The learner is the agent. Everything it interacts with is the environment. Our robot is the agent; the maze, its doors, and its rules are the environment.
At each step, the robot observes its situation, chooses an action, and receives a numerical reward. The environment changes, and the loop repeats. In our maze, every step costs −1, and reaching the exit adds a bonus of +10. Exiting in two steps therefore earns a total of 8. The objective is to learn behavior that earns a high total reward over time.
The essential difficulty: an action can be useful now and harmful later. RL must connect a later outcome to the earlier decisions that helped cause it. This is the credit assignment problem.
State and observation: the situation versus what you can see
A state, written s, contains the information needed to predict what happens next, given an action. In a simple maze, that might be the robot’s position. If locked doors and limited battery matter, position alone is insufficient: we also need the key inventory and battery level.
This is the Markov property: once we know the current state and action, older history adds no information about the next-state and reward distribution. A problem with this structure is a Markov decision process (MDP).
An observation, written o, is what the agent actually receives. A camera image may hide the key behind a wall. When observations do not reveal the full state, the problem is partially observable—a POMDP. Memory, a recurrent network, or a belief about hidden state can help. Merely calling a vector “the state” does not make it Markov.
Reward, return, and discount: three different things
A reward is one piece of feedback. The return, written G, adds up rewards from a point onward. A common objective discounts later rewards:
The discount factor γ (gamma) controls how heavily the objective weights the future. With γ = 0, only the next reward counts. A value near 1 gives distant consequences more weight. Finite episodes can use 1; continuing discounted problems commonly use a value below 1. Discounting changes the objective—it is not a prediction of whether a reward will arrive.
For the numerical maze examples below, we use γ = 1: simply add the rewards until the attempt ends.
A complete attempt, from starting the maze to reaching an ending, is an episode. A trajectory is the sequence of observations or states, actions, and rewards. A rollout is a collected sequence, which may stop before the episode ends. These words overlap, but “rollout” does not necessarily mean a complete game.
The reward must express what we actually want. If the robot gets a bonus every time it touches a key, it might learn to pick it up and drop it repeatedly. A high training reward is evidence of success only when the reward and evaluation capture the intended task.
2. Policy, value, model: three objects to remember
Most of the vocabulary becomes easier if we separate three things an agent might know.
Policy
What should I do?
A rule for choosing an action.Value
How good is this?
A prediction of future return.Model
What happens next?
A prediction of consequences.A policy is the decision rule
A policy, written π (pi), maps the agent’s information to an action or a distribution over actions. A deterministic policy might always choose right at this junction. A stochastic policy might assign left a probability of 30% and right 70%.
Sampling from that distribution means choosing right about 70% of the time across repeated choices in the same situation. Taking the most probable action means always choosing right here. Those are different decision rules; an implementation must specify which it uses.
Choosing the most probable action at deployment is common, but it changes a stochastic policy and is not always better. Some tasks benefit from retaining randomness.
A value is a forecast, not an immediate reward
The state-value function Vπ(s) predicts the expected return starting in state s and following policy π. The action-value function Qπ(s, a) predicts the expected return if we first take action a, then follow that policy.
At our junction, Q(s, left) = 4 and Q(s, right) = 7 would favor right. Those numbers are predicted returns, not probabilities. The superscript matters: a route can be promising for a skilled policy and poor for one that gets lost later. Q*, read “Q-star,” denotes values under optimal future behavior.
The advantage compares an action with the policy’s usual prospects in that state:
One junction, four connected ideas
Keep the policy and action values we just introduced. For a fixed policy, its state value is the average of its action values, weighted by how often it chooses each action.
| Action | Policy probability | Action value Q |
|---|---|---|
| Left | 30% | 4 |
| Right | 70% | 7 |
State value: V = 0.3 × 4 + 0.7 × 7 = 6.1.
Right’s advantage: 7 − 6.1 = +0.9.
Left’s advantage: 4 − 6.1 = −2.1.
Left can still lead to a positive return. It is worse than this policy’s average prospects at this junction. That is what its negative advantage means.
These are illustrative values. In practice, we estimate them from data; separately learned Q and V estimates need not agree perfectly.
A model predicts consequences
An environment model predicts the next state and reward, or their probability distributions. It may be a supplied simulator or something learned from experience. A model lets the robot ask what would happen if it moved right without actually moving right.
A neural network that predicts Q-values is a machine-learning model in everyday language. In the specific phrase model-based RL, however, “model” refers to a model of environment consequences. DQN has a neural network and is still called model-free.
Keep the questions separate: a policy chooses; a value evaluates; an environment model predicts. An algorithm can combine them.
Value-based methods build the decision around learned action values. Policy-based methods directly learn the decision rule. Actor–critic combines a learned policy with a value estimator. These are useful descriptions, not three completely separate boxes; “model-based” adds another dimension.
A map before the method names
Three other pairs of labels answer separate questions. We will meet them throughout the article:
- Model-free / model-based: do we learn values or a policy directly from experience, or also use a model of consequences to plan or imagine experience?
- On-policy / off-policy: must the experience come from the policy being updated, or can the update learn about a different policy from the one that collected it?
- Online / offline: can training collect more interaction, or must it work from a fixed dataset?
An agent can be model-free, off-policy, and online at the same time. The labels describe different aspects of it. As we encounter each method, keep asking: what is learned, how is it updated, and how does it choose?
For the formal definitions, see Spinning Up: Key Concepts in RL.
3. Learning a score for each action
Suppose we give the robot a notebook. Each row is a maze state; each column is an action. The entries estimate the return from taking that action. This is a Q-table.
To exploit what it knows, the robot selects the action with the highest Q-value. To discover alternatives, it must sometimes explore. In epsilon-greedy exploration, it takes a random available action with probability ε and otherwise takes a highest-valued action. A greedy policy is still a policy, even when there is no separate policy network.
Monte Carlo and temporal difference: when do we learn?
One option is to finish the attempt and use the actual return to update the notebook. This is a Monte Carlo (MC) target: learn from the sampled outcome. It needs no estimate of what lies beyond that completed return, but outcomes can be noisy and feedback arrives late.
A temporal-difference (TD) target combines the next observed reward with an estimate of what follows. For a state-value estimate, a one-step target is r + γV(next state). Updating an estimate using another estimate is called bootstrapping. It lets learning proceed before the episode ends, at the cost of depending on imperfect predictions.
At our junction, suppose the state-value estimate is 6.1. The robot takes a step, receives −1, and estimates a remaining return of 8 from the next state. With our γ = 1, the TD target is −1 + 8 = 7. The gap 7 − 6.1 = 0.9 is the TD error. A tabular update with learning rate 0.5 moves the estimate halfway toward the target: 6.1 + 0.5 × 0.9 = 6.55. This update uses a prediction of the future; it has not waited to see the actual return.
The recursive idea—return now equals the next reward plus future return—is expressed by Bellman equations. They relate values across states; they are not a separate neural-network architecture.
n-step returns sit between one-step TD and a complete Monte Carlo return: collect several actual rewards, then bootstrap from an estimate of what remains. Eligibility traces are a related mechanism for spreading updates across recently visited states or actions. These are ways to assign credit, not new choices of policy architecture.
Q-learning versus SARSA: which future action?
Q-learning uses the highest estimated Q-value at the next state:
It learns toward greedy future behavior even when the data-collecting robot explores. SARSA instead uses the next action the robot actually selected:
SARSA’s name comes from the sequence state, action, reward, state, action. If the robot keeps exploring near a dangerous shortcut, SARSA’s target reflects that exploratory behavior. Q-learning’s target assumes a greedy choice next. This is the core distinction, rather than “one is careful and the other is reckless.”
Q-learning is off-policy: the policy being learned about can differ from the one collecting experience. Ordinary SARSA is on-policy: it evaluates and improves the policy generating its actions. At a true terminal state, both targets have no future-value term.
When to use a table: when the state and action sets are small enough to revisit and store. A table does not infer a useful value for an unseen state. Grouping states into bins can help, but the grouping decides which differences the agent can no longer distinguish.
DQN: replace the notebook with a neural network
A Deep Q-Network (DQN) predicts action values from features or observations. In a common setup, one network receives the state and outputs a score for each discrete action. At decision time, compare those scores; during training, usually add exploration.
The network can share what it learns across similar inputs. It does not guarantee good decisions in unfamiliar situations. Standard DQN uses experience replay, a buffer of old transitions sampled for training, and a slowly updated target network that supplies more stable learning targets. These serve different purposes: replay reuses and mixes experience; the target network slows the movement of the prediction target.
How does the network learn? For each sampled transition, construct a Q-learning target from the observed reward and the target network’s estimate of the best next action. Train the current network to bring its prediction for the action actually taken closer to that target. During use, select a highest-scoring available action. This is a learned comparison, without simulating future moves at that moment.
DQN has several useful refinements. We will distinguish Double DQN, dueling networks, and Rainbow in the optional reference section.
Why choose this family? It is a natural fit when the available actions are discrete and manageable, and reusing expensive experience matters. Its difficulty is learning reliable values while its own predictions influence future targets. Standard DQN also becomes awkward when selecting the best action requires searching a huge or continuous action space.
Source: Human-level control through deep reinforcement learning (DQN).
4. Learning the decision rule directly
Instead of learning a score and then searching for its maximum, we can directly adjust the policy. At the junction, the network might output π(left | s) = 0.3 and π(right | s) = 0.7. The vertical bar means “given this state.”
Policy gradient methods update the policy’s parameters to increase expected return. Intuitively, they make actions associated with better outcomes more likely. The gradient tells the optimizer how changing the parameters changes the objective; it does not require differentiating through the real environment.
REINFORCE: learn from completed attempts
REINFORCE is a classic Monte Carlo policy-gradient method. Run the policy, observe returns, and use them to weight the policy-gradient update for each sampled action. A positive weight encourages a sampled action; a negative weight discourages it. Without a baseline, that weight is the return, which can be negative. With a baseline, the weight is return minus baseline.
This works without a critic. The problem is noisy credit assignment: the robot might turn right sensibly and then lose because of a much later mistake. A single failed attempt is weak evidence that turning right was wrong. A baseline helps compare outcomes with what was normally expected.
Why start here? REINFORCE is useful for understanding policy learning and for simple episodic problems. Long, noisy episodes make its gradient estimates difficult to learn from efficiently. Williams’s original paper develops the underlying update.
Actor–critic: learn an action rule and an evaluator
The actor is the policy. The critic is a learned value estimator that helps train it. These are roles: they can be separate networks, or share a network with different output heads.
One complete learning cycle
Return to the original snapshot: right has probability 70%, and the critic predicts a state value of 6.1. The following example waits for the completed return so we can see the two learning jobs separately.
- Act. Sample an action from the actor’s distribution. This time, the robot chooses right.
- Observe. Complete the attempt. Suppose its actual return from the junction is 8.
- Train the critic. Compare its old prediction, 6.1, with the return target, 8. Adjust the critic’s weights to reduce that prediction error, commonly using squared error. The critic learns from experience; it is not given the correct value in advance.
- Train the actor. Use
8 − 6.1 = +1.9, computed with the old prediction, to weight the policy update. This outcome encourages choosing right more often. - Repeat. Collect more attempts under the updated policy. Fit the critic to those outcomes and keep improving the actor.
Why +1.9 here, but +0.9 earlier? Earlier we compared the expected return for right, 7, with 6.1. This time we observed one return of 8. The sampled learning signal is noisy; averaging over outcomes connects it to the expected advantage.
With a critic, we can also use short TD-based targets instead of waiting until the end. This can reduce variance and speed feedback, while introducing dependence on the critic’s accuracy. The critic is useful, not logically necessary. It makes the actor’s learning signal more informative.
In the earlier one-step example, −1 + V(next state) = 7 could train the critic before the episode finishes, and the TD error 7 − 6.1 = +0.9 could provide the actor’s learning signal. At deployment, the actor can choose an action without this training calculation.
A3C means Asynchronous Advantage Actor–Critic: multiple workers collect experience and update shared parameters asynchronously. Advantage Actor–Critic (A2C) is the synchronous counterpart, combining workers’ data for coordinated updates. They belong to the same actor–critic family; their names mainly distinguish how collection and updates are organized. The A3C paper explains the asynchronous approach.
Generalized advantage estimation (GAE) blends TD errors across several future steps to estimate advantage. It provides a practical bias–variance tradeoff through a parameter called λ (lambda). GAE is an estimation technique used inside methods such as PPO, not an alternative complete agent. The GAE paper develops this connection.
Here, variance means how much an estimate fluctuates across sampled experience. Bias means systematic error in the estimate. Reducing noise by relying more on an imperfect critic can trade one problem for the other.
Why learn a policy if Q-values already tell us what to do?
For four directions, comparing four Q-values is easy. Now let the robot choose any steering angle and motor torque. Finding the action that maximizes a learned Q-function becomes an optimization problem at every decision. An actor can produce an action directly.
Policies also naturally represent randomized behavior and structured action distributions. These are useful reasons to learn one. They do not mean policy methods are only for continuous actions. They work with discrete actions too, and having continuous state features is a different issue from having continuous actions.
The cost is that policy learning can require a lot of fresh experience, depending on the algorithm. Some policy methods are on-policy; others deliberately reuse old data. We need the particular algorithm’s description before judging this tradeoff.
5. TRPO and PPO: improve without changing too much at once
A policy update changes which situations the robot will encounter. A large update based on yesterday’s experience can produce behavior for which that experience is no longer informative.
Trust Region Policy Optimization (TRPO) addresses this by approximately optimizing an improvement objective subject to a limit on how far the policy distribution changes. It measures that change with KL divergence, a measure of the difference between probability distributions. Its constrained update is more involved than ordinary gradient optimization. TRPO is the important predecessor here.
Proximal Policy Optimization (PPO) offers a simpler approach. In the commonly used PPO-Clip variant, the objective stops rewarding certain probability changes once they become too large.
In our example, right had probability 0.7 when data was collected, and its sampled advantage was +1.9. Raising its probability should help the objective. With a clipping parameter of 0.2, the beneficial contribution for that sample plateaus once the new-to-old probability ratio passes 1.2—here, a new probability of 0.7 × 1.2 = 0.84.
This is not a hard cap on the probability. Other samples and shared parameters can still move it further. Clipping discourages excessive updates; it does not guarantee that every update improves the policy.
A typical PPO implementation collects fresh rollouts, estimates advantages with a critic, then trains both components on that batch: the actor with the clipped policy objective, and the critic to predict return targets. It performs several minibatch passes before collecting more experience. PPO supports both discrete and continuous actions. The PPO paper presents the objective and experiments.
PPO at training time: actor chooses, environment responds, critic helps judge the outcome, clipped updates adjust the actor.
PPO at deployment: the actor produces the action distribution. No tree search is inherent in PPO, and the critic is usually unnecessary for choosing the action.
Why choose PPO? I would consider it when fresh simulation is affordable and a widely supported actor–critic baseline is useful. If collecting transitions is the main expense, an algorithm that reuses a replay buffer may be a better starting point. That is a reason to compare methods, not a promise that one will win.
6. DDPG, TD3, and SAC: actors that reuse experience
Return to the robot with steering and torque controls. We want an actor to output actions directly, but we also want to reuse past experience. These algorithms combine both.
DDPG and TD3: an actor guided by Q-values
Deep Deterministic Policy Gradient (DDPG) learns a deterministic actor and a Q-function critic. The critic learns from replayed transitions. The actor is adjusted toward actions that the critic scores highly. Exploration normally adds noise during data collection.
This is efficient in principle, but it can exploit errors in its own critic: an action may look excellent because its estimated value is wrong. Twin Delayed DDPG (TD3) addresses this with two critics, delayed actor updates, and smoothing of target actions. Its targets use the smaller critic estimate to limit overestimation.
Why choose TD3? It is a candidate for continuous control when replay and a deterministic actor suit the task. For a new baseline in this family, its corrections give a reason to prefer it over plain DDPG. DDPG and TD3 describe the details.
SAC: reward plus room to explore
Soft Actor–Critic (SAC) learns a stochastic actor and uses off-policy value learning. Its objective balances reward with entropy, a measure of how spread out the policy distribution is. This encourages the agent to keep alternatives available rather than becoming certain too early.
“Soft” refers to this entropy-regularized objective. It does not mean gentle actions or soft constraints. A temperature parameter controls the reward–entropy balance and can itself be adjusted during learning.
During training, SAC updates its critics from replayed transitions and trains the actor to balance their value estimates with entropy. During use, it can sample an action directly from the actor; a deterministic evaluation rule is also possible. Neither choice requires tree search.
Why choose SAC? It is a strong candidate for continuous control when environment interactions are costly and replay is valuable. The original formulation targets continuous actions; discrete variants also exist. SAC is an actor–critic method and off-policy, so “actor–critic” cannot be a synonym for “on-policy.” The original SAC paper explains the maximum-entropy objective.
7. Planning: think before moving
Until now, the robot’s learned values or policy have handled the decision directly. Suppose it also has a simulator. It can test hypothetical moves, inspect their consequences, and compare continuations before acting.
This is planning. Learning improves reusable knowledge from experience. Planning uses a model to compute what to do. They can work together, but neither word implies the other.
Dynamic programming, Dyna, and model predictive control
If a small environment’s dynamics are known, value iteration repeatedly applies Bellman optimality updates. Policy iteration alternates evaluating a policy and improving it. These are classic dynamic-programming planning methods; no trial-and-error data collection is necessary when the model is already supplied.
Dyna combines real transitions with simulated transitions from a learned model to update values. Model predictive control (MPC) plans a finite sequence of actions, executes the first, then replans from the new situation. MPC can use a known model or a learned one; it is not inherently an RL learning algorithm. The shared attraction is using predicted consequences. The shared risk, with a learned model, is planning around predictions that are wrong.
MCTS: spend computation on promising futures
Monte Carlo tree search (MCTS) grows a search tree through repeated simulations. One search iteration descends through candidate actions, expands a new node when appropriate, evaluates the resulting position, and backs that evaluation up along the path.
Selection balances actions with good estimated outcomes against actions that deserve more investigation. Basic variants may evaluate a leaf with a simulated continuation to the end. Neural variants can use a value estimate and stop much earlier.
How deep does it search? There is no universal depth. A time or simulation budget limits the work, and branches can reach different depths. A better value estimator can make short searches useful; more simulations do not require every branch to reach a terminal state.
These simulations happen before committing to one real move. After making that move and observing its result, the agent can search again from the new situation.
| Search run | First action | Return estimate | Visits so far left / right |
|---|---|---|---|
| 1 | Right | 6 | 0 / 1 |
| 2 | Left | 8 | 1 / 1 |
| 3 | Left | 8 | 2 / 1 |
| 4 | Right | 4 | 2 / 2 |
| 5 | Left | 7 | 3 / 2 |
Search stops → execute left once
Here, the final rule picks the most visited action.
In this trace, search first investigates each direction. It then favors left’s better results, while revisiting the less-explored right branch. The exact order depends on the search’s exploration rule; these numbers illustrate the bookkeeping, not a benchmark result.
What happened inside simulation 5? Start at the junction—the tree’s root. Follow left and then further simulated moves. Add a newly reached position to the tree—expansion—and evaluate that position, or leaf. Combine its estimated remaining return with rewards along the path. Here that gives 7 from the junction.
Backing up means updating the statistics on the traversed path. Left’s root visit count goes from 2 to 3, and its average becomes (8 + 8 + 7) / 3 ≈ 7.67. The robot has not physically tried five moves. Only after search finishes does it take the chosen real action.
AlphaZero: a network guides search, and search teaches the network
AlphaZero uses a policy head and a value head on one shared network. The policy supplies initial action probabilities, called priors; the value evaluates newly reached positions. At each search step, it selects an action using a score that combines the search’s current value estimate with an exploration bonus. That bonus grows with the policy prior and shrinks as the action is visited more often.
In training, self-play search produces a distribution over moves from root visit counts. The policy learns to match that search distribution; the value learns from the eventual game outcome. Early self-play moves are sampled to encourage exploration.
During evaluation, AlphaZero also runs MCTS and typically picks the most visited root move. The raw policy probability is not the final decision rule. Returning to our maze analogy, right could begin with a 70% prior, yet search could favor left after discovering better continuations. The tiny trace above illustrates this kind of reconsideration, rather than reproducing AlphaZero’s exact selection formula.
Using only the trained policy is a possible faster deployment choice, but it changes the system and must be evaluated separately. The AlphaZero paper describes the search-and-learning loop.
Two uses of “Monte Carlo”: Monte Carlo learning targets use sampled returns. Monte Carlo tree search is a planning procedure. Sharing those words does not make them the same algorithm.
Why choose search with learning? Consider it when hypothetical transitions are reliable and cheap enough, and improving each decision is worth extra runtime. Without a usable simulator, or under a strict action-latency budget, a policy or value network that acts directly may be more practical.
AlphaZero plans with known game rules. Its policy training target comes from search; it is not PPO with a second network.
MuZero and Dreamer: learning what to imagine
MuZero learns a representation and dynamics useful for predicting rewards, policies, and values inside search. Its internal states need not reconstruct the visible world. This lets it plan without being given the environment’s transition rules.
Dreamer learns a world model and trains an actor–critic using imagined trajectories. The actor can then choose actions without a fresh tree search at every decision. So “model-based” does not automatically mean “MCTS at deployment.” MuZero and Dreamer illustrate two different uses of learned models.
8. Labels that answer different questions
We can now describe an algorithm without forcing every term into one hierarchy. For example: PPO is a model-free, on-policy policy-optimization algorithm, usually implemented with an actor and critic. Each part adds different information.
On-policy versus off-policy: whose experience can an update use?
The behavior policy collects the data. The target policy is the policy being evaluated or improved. Off-policy methods allow these to differ. On-policy methods rely on experience from the policy being updated, allowing whatever limited reuse or correction the algorithm explicitly supports.
PPO performs several updates on a fresh batch using probability ratios, but that does not make it a general replay-buffer algorithm. DQN, TD3, and SAC are designed to learn from replayed experience collected under older behavior.
Online versus offline: can we collect more data?
Online RL can interact with the environment as learning proceeds. Offline RL learns from a fixed dataset, without collecting new experience during that training process. This is a different distinction from on-policy versus off-policy.
A DQN agent with a replay buffer can still be learning online. Conversely, giving DQN a fixed dataset does not automatically make it a reliable offline learner: it can overvalue unsupported actions and cannot try them to correct the mistake. Offline methods such as Conservative Q-Learning (CQL) explicitly address this kind of problem. See the offline RL tutorial and CQL paper.
Model-free versus model-based: do we use a model of consequences?
Model-free methods learn values or policies without using an environment model for planning or generating imagined training experience. Model-based methods use such a model. Training a model-free agent inside a simulator does not by itself make the learning algorithm model-based: interacting with the simulator for data differs from using it to reason over hypothetical futures.
Tabular versus deep: how is the knowledge represented?
A table stores separate entries. A function approximator shares parameters across inputs. Deep RL uses deep neural networks for one or more components; it does not identify one particular learning rule. Q-learning can be tabular, while DQN uses a network. Policies and environment models can also use networks.
Discrete versus continuous: which quantity are we discussing?
A robot can observe continuous positions and choose among four discrete moves. Another can observe the same positions and output continuous torque. The action space affects how actions are selected; the observation or state space affects how situations are represented. Neither is the same as continuing versus episodic, which asks whether the task has natural endings.
9. How to choose a method
Choose based on the problem’s constraints, then compare candidates under the same evaluation budget. The table gives starting points, not a ranking.
| Method | How it chooses | Why consider it? | Main cost or limitation |
|---|---|---|---|
| Tabular Q-learning / SARSA | Compare table entries; add exploration while learning. | Small, enumerable problems; transparent learning. | No automatic generalization to unseen states. |
| DQN family | Compare predicted Q-values for discrete actions. | Manageable action set; reuse experience. | Value errors can feed into targets; large action searches are awkward. |
| REINFORCE | Sample from a learned policy. | Understand policy gradients; simple episodic baseline. | Noisy returns and delayed feedback. |
| A2C / A3C | Actor selects; critic helps train it. | Direct actor–critic baseline with parallel experience collection. | Typically needs substantial fresh experience. |
| PPO / TRPO | Actor selects; constrain or discourage large updates. | Discrete or continuous actions; affordable fresh rollouts. | Limited old-data reuse; careful updates are not a success guarantee. |
| TD3 | Deterministic actor outputs continuous actions. | Continuous control with replay. | Actor depends on accurate critics and useful exploration. |
| SAC | Stochastic actor; reward–entropy objective. | Continuous control with replay and sustained exploration. | More components and an entropy tradeoff to manage. |
| AlphaZero-style search | Network-guided MCTS; root visit counts determine the move. | Reliable, affordable simulation; time to deliberate. | Many hypothetical transitions per actual decision. |
| Learned-model methods | Plan in a learned model, or train an actor in imagination. | Use predicted experience to reduce real interaction needs. | Model errors can produce misleading plans or training data. |
I would ask these questions before picking an acronym:
- Does an action affect later decisions? If each choice only has an immediate payoff and does not affect future opportunities, a bandit formulation may be enough. A contextual bandit also conditions the choice on current context. It avoids solving an unnecessary sequential problem.
- Can I afford to collect experience? Cheap parallel simulation makes PPO more plausible. Expensive interaction makes replay or model learning more attractive. A fixed historical dataset calls for offline-RL considerations.
- How hard is it to select an action? A few discrete options suit direct Q-value comparison. Continuous controls favor methods with actors. Enormous structured action spaces need more than simply choosing “discrete” in a library.
- Can I simulate alternatives at decision time? A reliable simulator and generous latency budget make search worth testing. A strict deadline can favor a single network evaluation.
- What does success cost? Measure sample efficiency—how many environment interactions are needed—separately from training compute, wall-clock time, and action latency. They are not interchangeable.
Use held-out situations and multiple random seeds, and compare with a simple baseline. Keep evaluation conditions consistent: a greedy DQN policy, a sampled PPO policy, and an AlphaZero agent spending seconds on search have different decision procedures and compute costs.
Also check the formulation before blaming the algorithm. Missing state information, misleading rewards, unavailable actions, or confusing a time-limit cutoff with a genuine terminal state can derail any method. A time cutoff may still require bootstrapping; a true ending does not.
10. Optional: further methods and names
You can skip to the glossary without losing the main story. This section places additional names you may encounter in papers.
DQN refinements: different fixes, not synonyms
- Double DQN separates selecting the best next action from evaluating that action in the learning target. This helps reduce overestimation.
- Dueling DQN changes the network architecture to combine a state-value component with relative action advantages. It shares information across actions. “Double” and “dueling” solve different problems and can be combined.
- C51 is a distributional DQN method: it represents returns using probabilities on 51 fixed return values, rather than predicting only one mean.
- Quantile Regression DQN (QR-DQN) also predicts a return distribution, but learns return values at fixed percentile levels. These are its quantiles. C51 fixes the values and learns their probabilities; QR-DQN fixes the probability levels and learns the values.
- Rainbow combines several DQN improvements, including Double DQN, a dueling architecture, and distributional learning. It names a combination, not a competing definition of Q-learning.
Sources: Double DQN, dueling networks, C51, QR-DQN, and Rainbow.
Other directions
These ideas extend the map rather than replacing it:
- Distributional RL predicts a distribution of returns instead of only their mean. This models variability in outcomes; it is not automatically a measure of uncertainty caused by limited knowledge. C51 and QR-DQN above are examples.
- Imitation learning learns behavior from demonstrations. Behavior cloning treats demonstrated actions as supervised targets. Inverse RL tries to infer a reward explaining behavior. Neither term is a synonym for offline RL.
- Hierarchical RL organizes decisions across timescales: choosing “go to the door” can invoke a skill that takes many lower-level steps. An option formalizes a temporally extended action.
- Multi-agent RL studies interacting learners. Cooperation, competition, and changing opponents introduce complications beyond a single fixed environment. Self-play is a training setup where agents learn by playing versions of themselves.
- RLHF, reinforcement learning from human feedback, describes the source of feedback rather than one optimizer. Human preferences can train a reward model, which then guides policy learning. PPO is one possible optimizer. Direct Preference Optimization (DPO) instead learns from preferred and rejected responses using a direct preference objective, without a separately trained reward model in its standard formulation. It is not another name for PPO.
Examples and primary sources: distributional RL, inverse RL, options, learning from human preferences, and DPO.
11. A glossary without the tangled names
These are the equivalences worth remembering—and the ones to avoid assuming.
- Policy · actor · policy network
- A policy is the decision rule. “Actor” names its role in actor–critic. A policy network is one way to represent it. A policy need not be a neural network.
- Value · V · Q · critic
- “Value function” is the umbrella. V evaluates a state; Q evaluates a state–action pair. A critic estimates values to help train a policy; it can estimate V or Q.
- Reward · return · value · advantage
- Reward: one piece of feedback. Return: accumulated future rewards, possibly discounted. Value: expected return. Advantage: how an action’s value compares with the policy’s state value.
- Transition · trajectory · rollout · episode
- A transition is one step, often stored as (state, action, reward, next state, terminal flag). A trajectory is a sequence. A rollout collects a sequence, possibly partial. An episode is a complete attempt between task boundaries.
- Policy gradient · policy optimization · actor–critic
- Policy optimization is the broader goal of improving a policy directly. Policy gradients use gradient estimates to do so. Actor–critic describes the combination of a policy with a learned evaluator. These labels overlap.
- Model · world model · model-free
- In this taxonomy, a model predicts environment consequences; “world model” emphasizes a learned one. Model-free does not mean “no neural network.” A reward model alone is not a complete dynamics model.
- Prediction · control
- Prediction evaluates behavior under a given policy. Control seeks a better policy. In RL, “control” does not require motors or continuous actions.
- Exploration · exploitation · entropy
- Exploration gathers information; exploitation uses current knowledge to pursue reward. Entropy measures a distribution’s spread. Randomness can help exploration, but is not the same as useful information gathering.
- Learning rate · discount factor
- The learning rate controls the size of an update. The discount factor controls the objective’s weighting of future rewards. They solve different problems.
- Bootstrapping · replay · target network
- Bootstrapping uses an estimate inside a learning target. Replay reuses stored experience. A target network supplies more slowly changing estimates. None is a synonym for another.
- Sparse reward · dense reward · reward shaping
- Sparse rewards arrive rarely; dense rewards provide more frequent feedback. Shaping adds feedback to aid learning. Unless carefully designed, it can change which behavior is optimal.
- Model-free/model-based · on-policy/off-policy · online/offline
- Three independent questions: do we use a model of consequences; whose behavior produced the data; and can training collect more interaction?
Back at the junction
The robot is still choosing between left and right. A Q-learning agent compares learned action values. A policy-gradient agent uses a learned action rule. An actor–critic agent learns that rule with help from a value estimator. A planning agent investigates possible futures before committing.
What differs is the machinery producing the decision: the objects being learned, the experience used to update them, and the computation spent before acting. Once those pieces are explicit, the acronyms become easier to place—and choosing a method becomes a question about the problem rather than its popularity.
Further reading
For a short refresher, begin with Spinning Up’s introduction, followed by its algorithm taxonomy and policy-gradient explanation. For a fuller foundation, use Sutton and Barto’s Reinforcement Learning: An Introduction. David Silver’s course provides lecture slides and worked explanations. The papers linked alongside the methods are useful once the basic distinctions feel familiar.