Retractions
A retraction is how every step in this package is taken. An OptimizerMethod produces a direction in the horizontal component $\mathfrak{g}^\mathrm{hor}$ of the Lie algebra, and the retraction turns that direction back into a point of the manifold. Two of them ship with the package — Cayley, the Cayley transform, and Geodesic, the exponential map — and since 0.2.0 Geodesic additionally carries an algorithm that says how the exponential is evaluated. There are five of those, and the choice between them is a numerical one: they compute the same map and differ in accuracy at a large step, in cost, and in which backends they run on.
This page is the retractions themselves: what a retraction is, how the two of them are derived on a homogeneous space, where one sits in the optimizer, and how Cayley and Geodesic differ as maps. The numerical problem of evaluating the exponential — the five algorithms, what they trade, and which one to reach for — is on the Exponential Algorithms page. The optimizer that uses either retraction is described on the Optimization on Homogeneous Spaces page.
What a retraction is
In practice we usually do not solve the geodesic equation exactly in each optimization step (even though this is possible and computationally feasible), but prefer approximations that are called "retractions" [8] for numerical stability. The definition of a retraction in GeometricOptimizers is slightly different from how it is usually defined in textbooks [8, 11]. We discuss these differences here.
Classical Retractions
By "classical retraction" we here mean the textbook definition.
A classical retraction is a smooth map
\[R: T\mathcal{M}\to\mathcal{M}:(x,v)\mapsto{}R_x(v),\]
such that each curve $c(t) := R_x(tv)$ is a local approximation of a geodesic, i.e. the following two conditions hold:
- $c(0) = x$ and
- $c'(0) = v.$
Perhaps the most common example for matrix manifolds is the Cayley retraction. It is a retraction for many matrix Lie groups [11–13].
The Cayley retraction for $V\in{}T_\mathbb{I}G\equiv\mathfrak{g}$ is defined as
\[\mathrm{Cayley}(V) = \left(\mathbb{I} - \frac{1}{2}V\right)^{-1}\left(\mathbb{I} +\frac{1}{2}V\right).\]
We show that the Cayley transform is a retraction for $G = SO(N)$ at $\mathbb{I}\in{}SO(N)$:
Proof
The Cayley transform trivially satisfies $\mathrm{Cayley}(\mathbb{O}) = \mathbb{I}$. So what we have to show is the second condition for a retraction and that $\mathrm{Cayley}(V)\in{}SO(N)$. For this take $V\in\mathfrak{so}(N).$ We then have
\[\frac{d}{dt}\bigg|_{t = 0}\mathrm{Cayley}(tV) = \frac{d}{dt}\bigg|_{t = 0}\left(\mathbb{I} - \frac{1}{2}tV\right)^{-1}\left(\mathbb{I} +\frac{1}{2}tV\right) = \frac{1}{2}V - \frac{1}{2}V^T = V,\]
which satisfies the second condition. We further have
\[\frac{d}{dt}\bigg|_{t = 0}(\mathrm{Cayley}(tV))^T\mathrm{Cayley}(tV) = (\frac{1}{2}V - \frac{1}{2}V^T)^T + \frac{1}{2}V - \frac{1}{2}V^T = 0.\]
This proves that the Cayley transform maps to $SO(N)$.
We should mention that the factor $\frac{1}{2}$ is sometimes left out in the definition of the Cayley transform when used in different contexts. But it is necessary for defining a retraction as without it the second condition is not satisfied.
We can also use the Cayley retraction at a different point than the identity $\mathbb{I}.$ For this consider $\bar{A}\in{}SO(N)$ and $\bar{B}\in{}T_{\bar{A}}SO(N) = \{\bar{B}\in\mathbb{R}^{N\times{}N}: \bar{A}^T\bar{B} + \bar{B}^T\bar{A} = \mathbb{O}\}$. We then have $\bar{A}^T\bar{B}\in\mathfrak{so}(N)$ and
\[ \overline{\mathrm{Cayley}}: T_{\bar{A}}SO(N) \to SO(N), \bar{B} \mapsto \bar{A}\mathrm{Cayley}(\bar{A}^T\bar{B}),\]
is a retraction $\forall{}\bar{A}\in{}SO(N)$.
As a retraction is always an approximation of the geodesic map, we now compare the cayley retraction for the example we introduced along Riemannian manifolds:
η_increments = 0.2 : 0.2 : 5.4
Δ_increments = [Δ * η for η in η_increments]
Y_increments_geodesic = [geodesic(Y, Δ_increment) for Δ_increment in Δ_increments]
Y_increments_cayley = [cayley(Y, Δ_increment) for Δ_increment in Δ_increments]

We see that for small $\Delta$ increments the Cayley retraction seems to match the geodesic retraction very well, but for larger values there is a notable discrepancy. We can plot this discrepancy directly:
zip_ob = zip(Y_increments_geodesic, Y_increments_cayley, axes(Y_increments_geodesic, 1))
discrepancies = [norm(Y_geo_inc - Y_cay_inc) for (Y_geo_inc, Y_cay_inc, _) in zip_ob]
nothing┌ Warning: `arrows` are deprecated in favor of `arrows2d` and `arrows3d`.
└ @ Makie ~/.julia/packages/Makie/frKXt/src/basic_recipes/arrows.jl:204

In GeometricOptimizers
The way we use retractions[1] in GeometricOptimizers is slightly different from their classical definition:
Given a section $\lambda:\mathcal{M}\to{}G,$ where $\mathcal{M}$ is a homogeneous space, a retraction is a map $\mathrm{Retraction}:\mathfrak{g}^\mathrm{hor}\to{}G$ such that
\[\Delta \mapsto \lambda(Y)\mathrm{Retraction}(\lambda(Y)^{-1}\Omega(\Delta)\lambda(Y))E,\]
is a classical retraction.
This map $\mathrm{Retraction}$ is also what was visualized in the figure on the general optimization framework. We now discuss how the geodesic retraction (exponential map) and the Cayley retraction are implemented in GeometricOptimizers.
Retractions for Homogeneous Spaces
Here we harness special properties of homogeneous spaces to obtain computationally efficient retractions for the Stiefel manifold and the Grassmann manifold. This is also discussed in e.g. [9, 12].
The geodesic retraction is a retraction whose associated curve is also the unique geodesic. For many matrix Lie groups (including $SO(N)$) geodesics are obtained by simply evaluating the exponential map [8, 14]:
The geodesic on a compact matrix Lie group $G$ with bi-invariant metric for $\bar{B}\in{}T_{\bar{A}}G$ is simply
\[\gamma(t) = \exp(t\cdot{}\bar{B}\bar{A}^{-1})\bar{A} = \bar{A}\exp(t\cdot{}\bar{A}^{-1}\bar{B}),\]
where $\exp:\mathfrak{g}\to{}G$ is the matrix exponential map.
The last equality in the equation above is a result of:
\[\begin{aligned} \exp(\bar{A}^{-1}\hat{B}\bar{A}) = \sum_{k=0}^\infty\frac{1}{k!}(\bar{A}^{-1}\hat{B}\bar{A})^k & = \sum_{k=0}^\infty \frac{1}{k!}\underbrace{(\bar{A}^{-1}\hat{B}\bar{A})\cdots(\bar{A}^{-1}\hat{B}\bar{A})}_{\text{$k$ times}} \\ & = \sum_{k=0}^\infty \frac{1}{k!} \bar{A}^{-1} \hat{B}^k \bar{A} = \bar{A}^{-1}\exp(\hat{B})\bar{A}. \end{aligned}\]
Because $SO(N)$ is compact and we furnish it with the canonical metric, i.e.
\[ g:T_{\bar{A}}G\times{}T_{\bar{A}}G \to \mathbb{R}, (B_1, B_2) \mapsto \mathrm{Tr}(B_1^TB_2) = \mathrm{Tr}((B_1\bar{A}^{-1})^T(B_2\bar{A}^{-1})),\]
its geodesics are thus equivalent to the exponential maps. We now use this observation to obtain an expression for the geodesics on the Stiefel manifold. We use the following theorem from [14, Proposition 25.7]:
The geodesics for a naturally reductive homogeneous space $\mathcal{M}$ starting at $Y$ are given by:
\[\gamma_{\Delta}(t) = \exp(t\cdot\Omega(\Delta))Y,\]
where the $\exp$ is the exponential map for the Lie group $G$ corresponding to $\mathcal{M}$.
The theorem requires the homogeneous space to be naturally reductive:
A homogeneous space is called naturally reductive if the following two conditions hold:
- $\bar{A}^{-1}\bar{B}\bar{A}\in\mathfrak{g}^\mathrm{hor}$ for every $\bar{B}\in\mathfrak{g}^\mathrm{hor}$ and $\bar{A}\in\exp(\mathfrak{g}^\mathrm{ver}$),
- $g([X, Y]^\mathrm{hor}, Z) = g(X, [Y, Z]^\mathrm{hor})$ for all $X, Y, Z \in \mathfrak{g}^\mathrm{hor}$,
where $[X, Y]^\mathrm{hor} = \Omega(XYE - YXE)$. If only the first condition holds the homogeneous space is called reductive (but not naturally reductive).
We state here without proof that the Stiefel manifold and the Grassmann manifold are naturally reductive. We can however provide empirical evidence here:
B̄ = rand(SkewSymMatrix, 6) # ∈ 𝔤
Ā = exp(B̄ - StiefelLieAlgHorMatrix(B̄, 3)) # ∈ exp(𝔤ᵛᵉʳ)
X = rand(StiefelLieAlgHorMatrix, 6, 3) # ∈ 𝔤ʰᵒʳ
Y = rand(StiefelLieAlgHorMatrix, 6, 3) # ∈ 𝔤ʰᵒʳ
Z = rand(StiefelLieAlgHorMatrix, 6, 3) # ∈ 𝔤ʰᵒʳ
Ā' * X * Ā # this has to be in 𝔤ʰᵒʳ for St(3, 6) to be reductive6×6 Matrix{Float64}:
0.0 -0.914247 -0.387374 -0.650933 -0.741877 -0.7678
0.914247 0.0 -0.771916 -0.454759 -0.869684 -0.80972
0.387374 0.771916 0.0 -0.924238 -0.292897 -0.749455
0.650933 0.454759 0.924238 0.0 0.0 0.0
0.741877 0.869684 0.292897 0.0 0.0 0.0
0.7678 0.80972 0.749455 0.0 0.0 0.0verifies the first property and
adʰᵒʳ(X, Y) = StiefelLieAlgHorMatrix(X * Y - Y * X, 3)
tr(adʰᵒʳ(X, Y)' * Z) ≈ tr(X' * adʰᵒʳ(Y, Z))trueverifies the second.
In GeometricOptimizers we always work with elements in $\mathfrak{g}^\mathrm{hor}$ and the Lie group $G$ is always $SO(N)$. We hence use:
\[ \gamma_\Delta(t) = \exp(\lambda(Y)\lambda(Y)^{-1}\Omega(\Delta)\lambda(Y)\lambda(Y)^{-1})Y = \lambda(Y)\exp(\lambda(Y)^{-1}\Omega(\Delta)\lambda(Y))E.\]
Based on this we define the maps:
\[\mathtt{geodesic}: \mathfrak{g}^\mathrm{hor} \to G, \bar{B} \mapsto \exp(\bar{B}),\]
and
\[\mathtt{cayley}: \mathfrak{g}^\mathrm{hor} \to G, \bar{B} \mapsto \mathrm{Cayley}(\bar{B}),\]
where $\bar{B} = \lambda(Y)^{-1}\Omega(\Delta)\lambda(Y)$. These expressions for geodesic and cayley are the ones that we typically use in GeometricOptimizers for computational reasons. We show how we can utilize the sparse structure of $\mathfrak{g}^\mathrm{hor}$ for computing the geodesic retraction and the Cayley retraction (i.e. the expressions $\exp(\bar{B})$ and $\mathrm{Cayley}(\bar{B})$ for $\bar{B}\in\mathfrak{g}^\mathrm{hor}$). Similar derivations can be found in [12, 15, 16].
Further note that, even though the global section $\lambda:\mathcal{M} \to G$ is not unique, the final geodesic $\gamma_\Delta(t) = \lambda(Y)\exp(\lambda(Y)^{-1}\Omega(\Delta)\lambda(Y))E$ does not depend on the particular section we choose.
The Geodesic Retraction
An element $\bar{B}$ of $\mathfrak{g}^\mathrm{hor}$ can be written as:
\[\bar{B} = \begin{bmatrix} A & -B^T \\ B & \mathbb{O} \end{bmatrix} = \begin{bmatrix} \frac{1}{2}A & \mathbb{I} \\ B & \mathbb{O} \end{bmatrix} \begin{bmatrix} \mathbb{I} & \mathbb{O} \\ \frac{1}{2}A & -B^T \end{bmatrix} =: B'(B'')^T,\]
where we exploit the sparse structure of the array, i.e. it is a multiplication of a $N\times2n$ with a $2n\times{}N$ matrix.
We further use the following:
\[ \begin{aligned} \exp(B'(B'')^T) & = \sum_{n=0}^\infty \frac{1}{n!} (B'(B'')^T)^n = \mathbb{I} + \sum_{n=1}^\infty \frac{1}{n!} B'((B'')^TB')^{n-1}(B'')^T \\ & = \mathbb{I} + B'\left( \sum_{n=1}^\infty \frac{1}{n!} ((B'')^TB')^{n-1} \right)B'' =: \mathbb{I} + B'\mathfrak{A}(B', B'')B'', \end{aligned}\]
where we defined $\mathfrak{A}(B', B'') := \sum_{n=1}^\infty \frac{1}{n!} ((B'')^TB')^{n-1}.$ Note that evaluating $\mathfrak{A}$ relies on computing products of small matrices of size $2n\times2n.$ We do this by relying on a simple Taylor expansion, implemented as GeometricOptimizers.𝔄 (see the GeometricOptimizers documentation for its docstring).
The final expression we obtain is:
\[\exp(\bar{B}) = \mathbb{I} + B' \mathfrak{A}(B', B'') (B'')^T\]
The Cayley Retraction
For the Cayley retraction we leverage the decomposition of $\bar{B} = B'(B'')^T\in\mathfrak{g}^\mathrm{hor}$ through the Sherman-Morrison-Woodbury formula:
\[(\mathbb{I} - \frac{1}{2}B'(B'')^T)^{-1} = \mathbb{I} + \frac{1}{2}B'(\mathbb{I} - \frac{1}{2}B'(B'')^T)^{-1}(B'')^T\]
So what we have to compute the inverse of:
\[\mathbb{I} - \frac{1}{2}\begin{bmatrix} \mathbb{I} & \mathbb{O} \\ \frac{1}{2}A & -B^T \end{bmatrix}\begin{bmatrix} \frac{1}{2}A & \mathbb{I} \\ B & \mathbb{O} \end{bmatrix} = \begin{bmatrix} \mathbb{I} - \frac{1}{4}A & - \frac{1}{2}\mathbb{I} \\ \frac{1}{2}B^TB - \frac{1}{8}A^2 & \mathbb{I} - \frac{1}{4}A \end{bmatrix}.\]
By leveraging the sparse structure of the matrices in $\mathfrak{g}^\mathrm{hor}$ we arrive at the following expression for the Cayley retraction (similar to the case of the geodesic retraction):
\[\mathrm{Cayley}(\bar{B}) = \mathbb{I} + \frac{1}{2} B' \left(\mathbb{I}_{2n} - \frac{1}{2} (B'')^T B'\right)^{-1} (B'')^T \left(\mathbb{I} + \frac{1}{2} \bar{B}\right),\]
where we have abbreviated $\mathbb{I} := \mathbb{I}_N.$ We conclude with a remark:
As mentioned previously the Lie group $SO(N)$, i.e. the one corresponding to the Stiefel manifold and the Grassmann manifold, has a bi-invariant Riemannian metric associated with it: $(B_1,B_2)\mapsto \mathrm{Tr}(B_1^TB_2)$. For other Lie groups (e.g. the symplectic group) the situation is slightly more difficult.
One of such Lie groups is the group of symplectic matrices [12]; for this group the expressions presented here are more complicated.
Where a retraction sits in the algorithm
The optimizer never moves a point of the manifold directly. It keeps a section $\Lambda^{(t)} \in G$ with $Y^{(t)} = \Lambda^{(t)}E$, and one step is
\[\Lambda^{(t+1)} \gets \Lambda^{(t)}\,\mathrm{retraction}(W^{(t)}), \qquad Y^{(t+1)} \gets \Lambda^{(t+1)}E,\]
with $W^{(t)} \in \mathfrak{g}^\mathrm{hor}$ the direction the method produced. So what a retraction has to return is an element of the group, not of the manifold, and the point follows from it. This is the extended retraction of [10]; update_section! performs the first line and apply_section the second. Because the retracted element is in $G$, a retracted point is on the manifold by construction, and check — which measures $\|Y^TY - \mathbb{I}\|$ — returns nothing but accumulated round-off.
An instance is passed to Optimizer as retraction = Cayley() or retraction = Geodesic(). The default is Cayley.
Both retractions factor the lift
A horizontal lift is sparse, and both retractions exploit that in the same way. For the Stiefel manifold, lift_factors writes
\[\bar{B} = \begin{bmatrix} A & -B^T \\ B & \mathbb{O} \end{bmatrix} = \begin{bmatrix} \tfrac{1}{2}A & \mathbb{I} \\ B & \mathbb{O} \end{bmatrix} \begin{bmatrix} \mathbb{I} & \mathbb{O} \\ \tfrac{1}{2}A & -B^T \end{bmatrix} =: B'(B'')^T,\]
with two $N\times{}2n$ factors — for a GrassmannLieAlgHorMatrix the same expression with $A \equiv \mathbb{O}$. Whatever matrix function a retraction needs is then evaluated on the $2n\times{}2n$ product
\[X := (B'')^TB',\]
which is small even when $N$ is large. The matrix function is therefore priced by the number of columns $n$ and not by the dimension $N$ of the ambient space; what is left of the retraction is the assembly around it, which is $O(N^2n)$ and not the $O(N^3)$ an $N\times{}N$ matrix function would cost. And, as the next sections show, the factorisation is also where the accuracy of the exponential is decided, because $X$ is a considerably worse-behaved matrix than $\bar{B}$ is.
Cayley and Geodesic
Cayley is the Cayley transform,
\[\mathrm{Cayley}(\bar{B}) = \left(\mathbb{I} - \tfrac{1}{2}\bar{B}\right)^{-1} \left(\mathbb{I} + \tfrac{1}{2}\bar{B}\right),\]
which maps a skew-symmetric matrix into $SO(N)$ exactly. cayley never forms the $N\times{}N$ inverse: with the factorisation above it inverts a $2n\times{}2n$ matrix instead. No matrix function is involved anywhere, only a solve, so there is no series to cancel and no step size at which the transform breaks down the way an unscaled series does. That is not the same as being insensitive to the size of the lift: check still climbs from $10^{-15}$ to $3\cdot10^{-13}$ over the sweep in Staying on the manifold, which is the largest drift of anything measured there other than TaylorSeries.
Geodesic is the exponential map,
\[\mathrm{Geodesic}(\bar{B}) = \exp(\bar{B}),\]
i.e. the true geodesic of the manifold [17]. The difference that matters to the rest of the package is that $\alpha \mapsto \exp(\alpha\bar{B})$ is a one-parameter subgroup,
\[\exp\left((\alpha + \beta)\bar{B}\right) = \exp(\alpha\bar{B})\exp(\beta\bar{B}),\]
and $\alpha \mapsto \mathrm{Cayley}(\alpha\bar{B})$ is not. Everything in sight is a rational function of $\bar{B}$, so it all commutes and the two products can be compared directly: with $t = \alpha/2$ and $s = \beta/2$,
\[\mathrm{Cayley}(\alpha\bar{B})\,\mathrm{Cayley}(\beta\bar{B}) = \big[\mathbb{I} + (t + s)\bar{B} + ts\bar{B}^2\big] \big[\mathbb{I} - (t + s)\bar{B} + ts\bar{B}^2\big]^{-1} ,\]
against $[\mathbb{I} + (t+s)\bar{B}][\mathbb{I} - (t+s)\bar{B}]^{-1}$ for $\mathrm{Cayley}((\alpha+\beta)\bar{B})$. The two agree only where $ts\bar{B}^2 = \mathbb{O}$, and the gap is not small: on a random $\mathfrak{g}^\mathrm{hor}$ element of $\operatorname{St}(6,3)$ with $\|\bar{B}\| = 2.99$ it is $1.28$ at $\alpha = \beta = 1$, where the same difference for $\exp$ is $8\times10^{-16}$.
So the generator of the curve's velocity turns with $\alpha$ instead of staying $\bar{B}$. That is what retraction_differential supplies, and with it trial_slope is the exact derivative of a line search's merit function under either retraction. Before 0.2.0 the slope was paired against $\bar{B}$ regardless, which made it first-order under Cayley — 8.9% off at $\alpha = 0.5$ and 36% at $\alpha = 1$ on the $\operatorname{St}(6,3)$ problem of Linesearches on Manifolds, which is where that measurement lives.
The natural reading of "Cayley is a retraction and Geodesic is the exponential map" is that the slope was wrong because the curve was approximate. That is not the mechanism, and the smallest counterexample separates them. Take $N = 2$ and $\bar{B} = J = \left(\begin{smallmatrix}0 & -1\\ 1 & 0\end{smallmatrix}\right)$. Then
\[\mathrm{Cayley}(\alpha{}J) = \exp\big(2\arctan(\tfrac{\alpha}{2})\,J\big)\]
to the last bit — the Cayley curve is the geodesic, exactly, with nothing approximate about it. It is traversed at a different speed:
| $\alpha$ | 0 | 0.5 | 1 | 2 | 4 |
|---|---|---|---|---|---|
angle turned, Cayley | 0 | 0.4900 | 0.9273 | 1.5708 | 2.2143 |
angle turned, Geodesic | 0 | 0.5 | 1 | 2 | 4 |
$d\theta/d\alpha$ for Cayley | 1 | 0.9412 | 0.8 | 0.5 | 0.2 |
and that last row is exactly $D(\alpha) = \bar{B}(\mathbb{I} - \frac{\alpha^2}{4}\bar{B}^2)^{-1}$ at $J^2 = -\mathbb{I}$, i.e. $J/(1 + \alpha^2/4)$. Pairing the gradient against $J$ rather than against $D(\alpha)$ therefore overstates $\varphi'$ by $1 + \alpha^2/4$ — 6.2% at $\alpha = 0.5$, 25% at $\alpha = 1$, 100% at $\alpha = 2$. A curve can be exactly right and still give the wrong $\varphi'$, because $\varphi'$ is a derivative with respect to the parameter.
That also says how to read the percentages above: to leading order the error is $\alpha^2\lambda^2/4$ for an eigenvalue $\pm{}i\lambda$ of $\bar{B}$, so it grows with the step and with the size of the lift, and a figure quoted without its problem means little. The same measurement on the $\operatorname{St}(3,1)$ sphere of manifold_linesearch_tests.jl gives 4.5%, 18%, 72% and 288% at $\alpha = 0.25, 0.5, 1, 2$.
This distinction mattered in an earlier line-search regression: large trial steps amplified the difference between the Cayley curve and the geodesic. The exact differential fixed the derivative, while DEFAULT_STEP_CEILING fixed the excessive extrapolation. With both changes in place, the regression no longer depends on which retraction is selected.
The cost balance depends on the dimensions. cayley finishes with a product of two $N\times{}N$ matrices, whereas geodesic assembles $\mathbb{I} + B'\mathfrak{A}(X)(B'')^T$ from a $2n\times{}2n$ matrix function. The benchmark under What they cost shows one representative comparison.
That $2n\times{}2n$ matrix function is the whole of what remains, and it is not a formula but a numerical problem: which approximation to evaluate, and what to do when the argument is large. The Exponential Algorithms page takes it from here — the identity that reduces the exponential to $\mathfrak{A}$, the four algorithms that evaluate it and the fifth that sidesteps it, what each of them measures, and how to choose between them. Everything from here on is about the retractions as maps rather than as computations.
The retractions on the two manifolds
AbstractRetraction, Geodesic, Cayley, geodesic, cayley and retraction. geodesic and cayley each have a method on an AbstractLieAlgHorMatrix — the efficient form this page derives, for both the Stiefel and the Grassmann lift — and one on a point together with a tangent vector, which is the classical retraction of the footnote above. Their docstrings are on the reference page, where every docstring in the package is rendered once; the names above link to them.
Reference
- [8]
- P.-A. Absil, R. Mahony and R. Sepulchre. Optimization algorithms on matrix manifolds (Princeton University Press, Princeton, New Jersey, 2008).
- [9]
- T. Bendokat, R. Zimmermann and P.-A. Absil. A Grassmann manifold handbook: Basic geometry and computational aspects, arXiv preprint arXiv:2011.13699 (2020).
- [10]
- B. Brantner. Generalizing Adam To Manifolds For Efficiently Training Transformers, arXiv preprint arXiv:2305.16901 (2023).
- [11]
- E. Hairer, C. Lubich and G. Wanner. Geometric Numerical integration: structure-preserving algorithms for ordinary differential equations (Springer, Heidelberg, 2006).
- [12]
- T. Bendokat and R. Zimmermann. The real symplectic Stiefel and Grassmann manifolds: metrics, geodesics and applications, arXiv preprint arXiv:2108.12447 (2021).
- [13]
- B. Gao, N. T. Son, P.-A. Absil and T. Stykel. Riemannian optimization on the symplectic Stiefel manifold. SIAM Journal on Optimization 31, 1546–1575 (2021).
- [14]
- B. O'neill. Semi-Riemannian geometry with applications to relativity (Academic press, New York City, New York, 1983).
- [15]
- E. Celledoni and A. Iserles. Approximating the exponential from a Lie algebra to a Lie group. Mathematics of Computation 69, 1457–1480 (2000).
- [16]
- C. Fraikin, K. Hüper and P. V. Dooren. Optimization over the Stiefel manifold. In: PAMM: Proceedings in Applied Mathematics and Mechanics, Vol. 7 no. 1 (Wiley Online Library, 2007); pp. 1062205–1062206.
- [17]
- A. Edelman, T. A. Arias and S. T. Smith. The geometry of algorithms with orthogonality constraints. SIAM Journal on Matrix Analysis and Applications 20, 303–353 (1998).
- 1Classical retractions are also defined in
GeometricOptimizersunder the same name, i.e. there is e.g. a methodcayley(::StiefelLieAlgHorMatrix)and a methodcayley(::StiefelManifold, ::AbstractMatrix)(the latter being the classical retraction); but the user is strongly discouraged from using classical retractions as these are computationally inefficient.