Guided Policy Search using Sequential Convex Programming for Initialization of Trajectory Optimization Algorithms

Kim, Taewan; Elango, Purnanand; Malyuta, Danylo; Acikmese, Behcet

doi:10.23919/ACC53348.2022.9867151

Mathematics > Optimization and Control

arXiv:2110.06975 (math)

[Submitted on 13 Oct 2021 (v1), last revised 19 May 2022 (this version, v2)]

Title:Guided Policy Search using Sequential Convex Programming for Initialization of Trajectory Optimization Algorithms

Authors:Taewan Kim, Purnanand Elango, Danylo Malyuta, Behcet Acikmese

View PDF

Abstract:Nonlinear trajectory optimization algorithms have been developed to handle optimal control problems with nonlinear dynamics and nonconvex constraints in trajectory planning. The performance and computational efficiency of many trajectory optimization methods are sensitive to the initial guess, i.e., the trajectory guess needed by the recursive trajectory optimization algorithm. Motivated by this observation, we tackle the initialization problem for trajectory optimization via policy optimization. To optimize a policy, we propose a guided policy search method that has two key components: i) Trajectory update; ii) Policy update. The trajectory update involves offline solutions of a large number of trajectory optimization problems from different initial states via Sequential Convex Programming (SCP). Here we take a single SCP step to generate the trajectory iterate for each problem. In conjunction with these iterates, we also generate additional trajectories around each iterate via a feedback control law. Then all these trajectories are used by a stochastic gradient descent algorithm to update the neural network policy, i.e., the policy update step. As a result, the trained policy makes it possible to generate trajectory candidates that are close to the optimality and feasibility and that provide excellent initial guesses for the trajectory optimization methods. We validate the proposed method via a real-world 6-degree-of-freedom powered descent guidance problem for a reusable rocket.

Comments:	Presented in American Control Conference (ACC) 2022
Subjects:	Optimization and Control (math.OC); Robotics (cs.RO)
Cite as:	arXiv:2110.06975 [math.OC]
	(or arXiv:2110.06975v2 [math.OC] for this version)
	https://doi.org/10.48550/arXiv.2110.06975
Related DOI:	https://doi.org/10.23919/ACC53348.2022.9867151

Submission history

From: Taewan Kim [view email]
[v1] Wed, 13 Oct 2021 18:37:54 UTC (3,211 KB)
[v2] Thu, 19 May 2022 21:34:51 UTC (3,211 KB)

Mathematics > Optimization and Control

Title:Guided Policy Search using Sequential Convex Programming for Initialization of Trajectory Optimization Algorithms

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Mathematics > Optimization and Control

Title:Guided Policy Search using Sequential Convex Programming for Initialization of Trajectory Optimization Algorithms

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators