Computing Library › Reinforcement Learning
Reinforcement Learning

Continuous Control

Real-valued action spaces demand methods that output continuous commands rather than choosing from a discrete menu.

Actions that are numbers

Continuous control refers to reinforcement learning problems where actions are real-valued — a torque, a voltage, a coil current, a valve setting — rather than a discrete choice. Most physical control tasks are of this kind, and they require methods built for continuous action spaces, since value-based approaches that maximize over actions cannot enumerate an infinite set.

Why discrete methods struggle

Q-learning and DQN pick the action with the highest Q value by comparing all actions. With continuous actions that maximization is itself an optimization problem at every step. Discretizing the action space is possible but scales badly: fine control needs many bins, and multiple action dimensions multiply the count exponentially.

Methods built for it

Practical concerns

Continuous control raises issues discrete tasks avoid: actions must be bounded to physical limits (often via a squashing function), control effort should usually be penalized for smoothness, and small errors accumulate over long horizons. Exploration means perturbing continuous commands rather than sampling from a menu, which interacts with actuator limits and safety.

Where it applies

Robotics, vehicle control, and the control of physical plants are all continuous-control problems. Regulating a quantity toward a target with real-valued actuators — the shape of a task like holding a simulated plasma near a setpoint with continuously adjustable coil currents — falls squarely in this category and would use SAC, TD3, or PPO rather than any discrete-action method.