Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Guided policy search for a lightweight industrial robot arm
Luleå University of Technology, Department of Computer Science, Electrical and Space Engineering, Space Technology. Aalto University/Erasmus+.
2018 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

General autonomy is at the forefront of robotic research and practice. Earlier research has enabled robots to learn movement and manipulation within the context of a specific instance of a task and to learn from large quantities of empirical data and known dynamics. Reinforcement learning (RL) tackles generalisation, whereby a robot may be relied upon to perform its task with acceptable speed and fidelity in multiple---even arbitrary---task configurations. Recent research has advanced approximate policy search methods of RL, in which a function approximator is used to represent an optimal policy while avoiding calculation across the large dimensions of the state and action spaces of real robots. This thesis details the implementation and testing, on a lightweight industrial robot arm, of guided policy search (GPS), an RL algorithm that seeks to avoid the typical need, in machine learning, for lots of empirical behavioural samples, while maximising learning speed. GPS comprises a local optimal policy generator, here based on a linear-quadratic regulator, and an approximate general policy representation, here a feedforward neural network. A controller is written to interface an existing back-end implementation of GPS and the robot itself. Experimental results show that the GPS agent is able to perform basic reaching tasks across its configuration space with approximately 15 minutes of training, but that the local policies generated fail to be fully optimised within that timescale and that post-training operation suffers from oscillatory actions under perturbed initial joint positions. Further work is discussed and recommended for better training of GPS agents and making locally optimal policies more robust to disturbance while in operation.

Place, publisher, year, edition, pages
2018. , p. 65
National Category
Robotics and automation
Identifiers
URN: urn:nbn:se:ltu:diva-71217OAI: oai:DiVA.org:ltu-71217DiVA, id: diva2:1256071
Subject / course
Student thesis, at least 30 credits
Educational program
Space Engineering, master's level (120 credits)
Supervisors
Examiners
Available from: 2018-11-22 Created: 2018-10-16 Last updated: 2025-10-22Bibliographically approved

Open Access in DiVA

fulltext(15131 kB)630 downloads
File information
File name FULLTEXT01.pdfFile size 15131 kBChecksum SHA-512
18af70aa12992edd6b2766463e87d14c460030185b52f1cb63aa9166ebae33d34471773e2ffdd64682b2dc1a113ac4e10cf3097bf9885f1d8838a04b33fdab14
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
White, Jack
By organisation
Space Technology
Robotics and automation

Search outside of DiVA

GoogleGoogle Scholar
Total: 631 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 555 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf