Hyperspherical Contrastive Representations
Unlock Planning and Exploration

1 Seoul National University 2 Princeton University

* Equal contribution

“Can an RL agent discover how to stack five cubes with a gripper
from scratch, without prior data, rewards, demonstrations, or pretrained models?”

5-tower (gripper) BuilderBench

The sphere shows a spherical PCA projection (Liu et al., 2019) of the learned representations.

  • Current state ψ(s)
  • Goal ψ(g)
  • Inferred waypoints z1:K
  • Selected subgoal

TL;DR

The key is adaptive planning. Our method, C-Slerp, composes hyperspherical contrastive representations to decide how many waypoints each state-goal pair needs, and infers them in closed form via spherical linear interpolation (SLERP), with no high-level policy or graph search.


C-Slerp: Contrastive Planning via Spherical Linear Interpolation

Given a state and a goal , we use intermediate waypoints to decompose the long-horizon task into shorter subproblems. The four steps below show why spherical interpolation of learned contrastive representations gives these waypoints.

Hyperspherical contrastive representations

We parameterize temporal relationships using contrastive representations that lie on the unit hypersphere. We learn a state representation , a state-action representation , and a rotation matrix , and score temporal relationships with

We collect data online under a hierarchical goal-conditioned policy , where every steps a planner provides a waypoint representation for to pursue.
Using samples from this data-generating process, the representations are trained with contrastive learning on pairs separated by an interval drawn from . Under the symmetric InfoNCE objective, the Bayes-optimal takes the form

The learned rotation matrix can capture asymmetries in the dynamics, which a standard inner product cannot, since it is symmetric. This allows us to score an entire waypoint sequence by repeatedly applying a single to consecutive waypoint pairs as a reusable building block.

Planning as probabilistic inference

We take a probabilistic approach to planning. Given a state and goal , we infer intermediate waypoints along trajectories generated by , setting and . Consecutive waypoints and are separated by an interval , drawn independently for .
We approximate the resulting posterior with a tractable distribution :

Each segment from to spans one interval drawn from a geometric distribution, which is exactly the temporal relationship is trained to score. Scoring every segment with thus tells us how well a sequence of waypoints connects to .
For now, we fix and derive the corresponding optimal waypoints; we then derive a criterion for adaptively selecting for each state-goal pair in Geometric interpretation below (step 1, Select ).

Tractable reformulation

We factorize , with one distribution per waypoint , so that can be optimized in the same way for varying numbers of waypoints . Under the Bayes-optimal and the Markov assumption, the original KL objective becomes

It maximizes the expected sum of between waypoints while keeping each close to the marginal of visited states. Since the policy depends on a waypoint only through its representation , we infer instead of raw states, and the uniformity of the learned representations on turns the KL term above into an entropy term. The objective is then reformulated as follows (Appendix A.2):

This objective links the current state to the goal by repeatedly aligning consecutive representations through , while the entropy term encourages stochasticity.

Closed-form solution

We parameterize each as the distribution of a Gaussian sample normalized onto the unit hypersphere:

We weight the entropy term by . Under a small-variance approximation, the optimum of is given as follows (Appendix A.3):

where . When , the means lie evenly along the geodesic between and . In contrast, C-Slerp accounts for temporal transitions encoded in by interpolating from toward the rotated goal and mapping each interpolated point back using (steps 2–3 below). These closed-form expressions infer waypoint representations directly from the learned contrastive representations, without training an additional neural network.

Geometric interpretation of C-Slerp


Experimental Results

5-tower (gripper) BuilderBench

The sphere shows a spherical PCA projection (Liu et al., 2019) of the learned representations.

  • Current state ψ(s)
  • Goal ψ(g)
  • Inferred waypoints z1:K
  • Selected subgoal

How good is C-Slerp?

Learning curves on six BuilderBench manipulation tasks and six JaxGCRL locomotion tasks, comparing C-Slerp with CRL, RIS, Online HIQL, L3P, and SAC plus HER

Additional Experiments

See the paper (Appendix B) for detailed settings and further analysis.

What makes C-Slerp work?

Ablation on the number of waypoints K on 8-tower and Ant U5 Maze
Ablation on the selected waypoint index on 8-tower and Ant U5 Maze
  • Fixing the number of waypoints or the waypoint index that the policy pursues substantially degrades performance, and adding more waypoints alone does not help.
  • Too few waypoints make each subproblem too hard; too many slow progress with unnecessary stages. The plan should match the difficulty of the problem.
Single-goal Ant Big Maze and Ant U4 Maze environments
Learning curves on single-goal Ant Big Maze and Ant U4 Maze
  • With a single goal fixed throughout training and evaluation, C-Slerp achieves higher success rates than all baselines, while all baselines except CRL and L3P achieve zero success.
  • C-Slerp facilitates directed exploration toward distant goals.
Ablation on learning the rotation matrix R on 3-tower (gripper) and Ant U4 Maze
  • We compare learning with three fixed alternatives: the identity , and two random rotations whose Cayley parameters are sampled with .
  • Setting makes symmetric and prevents successful learning in both tasks. Random fixed rotations introduce asymmetry but yield unstable performance.
  • Asymmetry alone is insufficient; learning is necessary to capture temporal relationships between states.
CRL with larger discount factors compared with C-Slerp on 3-tower (gripper) and Ant U5 Maze
  • A natural way to adapt CRL to long-horizon tasks is to increase , since future states are sampled at offsets drawn from .
  • Across , CRL achieves low success rates, and increasing does not close the gap with C-Slerp.
  • Simply extending the temporal horizon of contrastive learning does not offer the same benefits as planning with an appropriate number of waypoints.

How do hyperparameters affect C-Slerp?

Ablation on network depth on 8-tower and Ant U5 Maze
  • Since C-Slerp uses contrastive representations, we examine whether it benefits from increased network depth, as reported by Wang et al. (2025).
  • Deeper networks generally achieve higher success rates. On the 8-tower task, depths of 16 and 32 perform comparably and outperform shallower networks.
Ablation on representation dimension on 8-tower and Ant U5 Maze
  • Performance is relatively robust once is sufficiently large, with further increases providing no consistent benefit.
  • On the 8-tower task, only achieves nonzero success. We hypothesize that this limited benefit arises because the encoder already provides sufficient nonlinear expressivity for the linear operator to model temporal relationships effectively.
Ablation on the threshold epsilon on 3-tower (gripper), 8-tower, Ant U4 Maze and Ant U5 Maze
  • A larger favors earlier waypoints, leading to more conservative subgoal selection.
  • C-Slerp performs well across in all evaluated environments except the 8-tower task, which is more sensitive to the threshold.
  • We use as the default configuration across all tasks.

Visualizations

Score colour scale from 0 to 1
Inferred waypoints and score heatmaps in Ant Big Maze Inferred waypoints and score heatmaps in Ant U5 Maze Inferred waypoints and score heatmaps in Ant Extreme Maze
  • Candidate waypoint
  • Selected waypoint
  • Goal
  • Current state
  • Goal
  • Candidate waypoint
  • Selected waypoint
  • Current state
  • Goal
  • We encode all positions across the entire maze, including walls, on a grid with 0.2 m spacing. For each inferred waypoint , we retrieve its nearest neighbor among these representations, and display at every grid point as a heatmap.
  • Although the waypoints are generated entirely in the latent space, they correspond to meaningful intermediate locations in the environment.
Evolution of K and i over time in JaxGCRL environments Evolution of K and i over time in BuilderBench environments
  • We plot the mean and sampled values of , the total number of transition segments to the goal, together with the selected waypoint index .
  • The mean of decreases as the agent approaches the goal, suggesting that the learned representation captures the progress of the agent toward the goal.

Final Remarks