Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Optimal hardware selection for AI-model learning
Luleå University of Technology, Department of Computer Science, Electrical and Space Engineering.
2025 (English)Independent thesis Advanced level (professional degree), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Artificial Intelligence (AI) is a field of study that has seen rapid advancements in recent times, particularly due to the development of larger and more complicated models. The scale of recent models, in particular Large Language Models (LLMs), has led to a massive increase in computation required for development of these state-of-the-art models. Out of necessity the workload centered around these large-scale models is therefore oftentimes delegated to purpose-built computational clusters. These clusters are usually heterogenous when it comes to the types of deployed hardware which leads to increased difficulty in scheduling incoming jobs. This thesis was carried out in collaboration with Ericsson AB with the goal of developing a recommendation system for research clusters. Specifically, the system should be able to make accurate predictions about key running time metrics of AI training jobs using only the source code as input, which could help improve cluster efficiency. To this end, the system was broken down into several sequential components which are able to operate independently. The final system consists of a Transformer-CRF model for feature extraction of the source code followed by a Set-Transformer based architecture for regression of the features into a set of key metrics. These models are combined with algorithmic and mathematical solutions to form the purpose-built system. The mixed results on practical data will be put in context of the available data at hand along with training time constraints, forming a framework to carry the system into a production-ready state.

Place, publisher, year, edition, pages
2025. , p. 59
Keywords [en]
machine learning, artificial intelligence, transformer, conditional random fields, code analysis, running time estimation, regression
National Category
Artificial Intelligence
Identifiers
URN: urn:nbn:se:ltu:diva-114430OAI: oai:DiVA.org:ltu-114430DiVA, id: diva2:1991789
External cooperation
Ericsson AB
Educational program
Computer Science and Engineering, master's level
Supervisors
Examiners
Available from: 2025-08-27 Created: 2025-08-25 Last updated: 2025-10-21Bibliographically approved

Open Access in DiVA

fulltext(839 kB)74 downloads
File information
File name FULLTEXT02.pdfFile size 839 kBChecksum SHA-512
3bb091551e753bc7a9531a5d2c2cdc48aedbb6dade44a0a5cce1cbf2a0669e230db3cc8f09e681d281dacdb4a8f80342f565e2e5692d711cb232f4f1471c4f27
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Furhoff, Hannes
By organisation
Department of Computer Science, Electrical and Space Engineering
Artificial Intelligence

Search outside of DiVA

GoogleGoogle Scholar
Total: 74 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 137 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf