Change search
Link to record
Permanent link

Direct link
Publications (10 of 190) Show all publications
Al-azzawi, S. S., Barney, E., Kovács, G. & Liwicki, M. (2027). A human-in-the-loop label error detection framework applied to Arabic-script HTR datasets. Pattern Recognition, 181, Article ID 114621.
Open this publication in new window or tab >>A human-in-the-loop label error detection framework applied to Arabic-script HTR datasets
2027 (English)In: Pattern Recognition, ISSN 0031-3203, E-ISSN 1873-5142, Vol. 181, article id 114621Article in journal (Refereed) Published
Abstract [en]

Despite recent advances, Handwritten Text Recognition (HTR) for Arabic-script languages still lags behind Latin-script HTR. Part of the problem is dataset quality. To help closing this gap, we propose a two-stage framework (CER-HV) for detecting label errors. Stage 1 (CER) is a Character-Error-Rate-based noise detector built on a Convolutional Recurrent Neural Network (CRNN) architecture. Stage 2 (HV) is the Human-In-The-Loop (HITL) Verification of noisy samples detected by the first stage. Applying the CER-HV framework on multiple Arabic-script datasets can identify samples with label errors including transcription, segmentation, orientation, and non-text content errors that can markedly affect HTR performance. These errors were identified by the first stage of the framework with up to 90% (top-50) precision.We also show that our CRNN achieves state-of-the-art performance across five of the six evaluated datasets, reaching 8.46% Character Error Rate (CER) on KHATT (Arabic), 8.22% on PHTI (Pashto), 10.59% on Ajami, and 10.11% on Muharaf (Arabic), all without any data cleaning. We establish a new baseline of 11.3% CER on the PHTD (Persian) dataset. Applying CER-HV improves evaluation CER by up to 1.8 percentage points after dataset cleaning and retraining. Although our experiments focus on documents written in an Arabic-script language, the framework is general and can be applied to other text recognition datasets.

Place, publisher, year, edition, pages
Elsevier Ltd, 2027
Keywords
Handwritten text recognition, Label error detection, CRNN, Pashto, Urdu, Persian, Ajami
National Category
Computer Sciences Natural Language Processing
Research subject
Machine Learning
Identifiers
urn:nbn:se:ltu:diva-119619 (URN)10.1016/j.patcog.2026.114621 (DOI)001856887300001 ()2-s2.0-105048023861 (Scopus ID)
Projects
MARTINA (20367152)
Funder
European Regional Development Fund (ERDF)
Note

Fulltext license: CC BY

Available from: 2026-09-02 Created: 2026-09-02 Last updated: 2026-09-02Bibliographically approved
Chintalapati, L. V., Jafari, A., Laufer, R., Liwicki, M. & Eickhoff, J. (2026). Advancing multi-class object detection from LEO/VLEO: model evaluation and onboard deployment tailored for a 16U CubeSat. CEAS Space Journal, 18, 417-444
Open this publication in new window or tab >>Advancing multi-class object detection from LEO/VLEO: model evaluation and onboard deployment tailored for a 16U CubeSat
Show others...
2026 (English)In: CEAS Space Journal, ISSN 1868-2502, E-ISSN 1868-2510, Vol. 18, p. 417-444Article in journal (Refereed) Published
Abstract [en]

On-board AI-based image recognition for Very Low Earth Orbit (VLEO) and Low Earth Orbit (LEO) missions is crucial for enabling timely Earth Observation while minimizing data transmission. This paper presents a proof-of-concept study for autonomous object detection, including the creation of a novel EO dataset and deployment on a 16U CubeSat's target hardware. The custom IceBrain-EO dataset contains 25,675 images across 12 classes. In this study, two models were benchmarked, with the fine-tuned YOLOv11m model achieving a mean Average Precision (mAP@.50) of 0.843. On the target NVIDIA Jetson Orin Nano edge device, this model demonstrated real-time performance, achieving an inference speed of 11.94 FPS on 640 × 640 images while consuming just 2.6W in its 15W power mode. These results validate the feasibility of autonomous payload operations and demonstrate a viable pathway for significantly reducing data downlink for future EO missions.

Place, publisher, year, edition, pages
Springer Nature, 2026
Keywords
 VLEO, Earth observation, EO dataset, Deep learning, Object detection, Edge deployment
National Category
Vehicle and Aerospace Engineering Computer graphics and computer vision
Research subject
Machine Learning; Space Systems
Identifiers
urn:nbn:se:ltu:diva-116376 (URN)10.1007/s12567-025-00697-6 (DOI)001671662400001 ()2-s2.0-105028881680 (Scopus ID)
Note

Full text license: CC BY

Available from: 2026-02-10 Created: 2026-02-10 Last updated: 2026-08-25Bibliographically approved
Brännvall, R., Zhang, T., Forsgren, H., Stoian, A., Sandin, F. & Liwicki, M. (2026). Inhibitor Transformers and Gated RNNs for Torus Efficient Fully Homomorphic Encryption.
Open this publication in new window or tab >>Inhibitor Transformers and Gated RNNs for Torus Efficient Fully Homomorphic Encryption
Show others...
2026 (English)Manuscript (preprint) (Other academic)
National Category
Computer Sciences
Research subject
Machine Learning
Identifiers
urn:nbn:se:ltu:diva-117524 (URN)
Available from: 2026-05-17 Created: 2026-05-17 Last updated: 2026-06-15
Shehzadi, T., Noor, R., Ifza, I., Liwicki, M., Stricker, D. & Afzal, M. Z. (2026). Knowledge Distillation in Object Detection: A Survey from CNN to Transformer. Sensors, 26, Article ID 292.
Open this publication in new window or tab >>Knowledge Distillation in Object Detection: A Survey from CNN to Transformer
Show others...
2026 (English)In: Sensors, E-ISSN 1424-8220, Vol. 26, article id 292Article, review/survey (Refereed) Published
Abstract [en]

Deep learning models, especially for object detection have gained immense popularity in computer vision. These models have demonstrated remarkable accuracy and performance, driving advancements across various applications. However, the high computational complexity and large storage requirements of state-of-the-art object detection models pose significant challenges for deployment on resource-constrained devices like mobile phones and embedded systems. Knowledge Distillation (KD) has emerged as a prominent solution to these challenges, effectively compressing large, complex teacher models into smaller, efficient student models. This technique maintains good accuracy while significantly reducing model size and computational demands, making object detection models more practical for real-world applications. This survey provides a comprehensive review of KD-based object detection models developed in recent years. It offers an in-depth analysis of existing techniques, highlighting their novelty and limitations, and explores future research directions. The survey covers the different distillation algorithms used in object detection. It also examines extended applications of knowledge distillation in object detection, such as improvements for lightweight models, addressing catastrophic forgetting in incremental learning, and enhancing small object detection. Furthermore, the survey also delves into the application of knowledge distillation in other domains such as image classification, semantic segmentation, 3D reconstruction, and document analysis.

Place, publisher, year, edition, pages
MDPI, 2026
Keywords
transformer, knowledge distillation, DETR, computer vision, deep neural networks
National Category
Computer Sciences Artificial Intelligence
Research subject
Machine Learning
Identifiers
urn:nbn:se:ltu:diva-116088 (URN)10.3390/s26010292 (DOI)001657639000001 ()41516728 (PubMedID)2-s2.0-105027110348 (Scopus ID)
Note

Full text license: CC BY 4.0;

Funder: AIRISE (101092312)

Available from: 2026-01-21 Created: 2026-01-21 Last updated: 2026-06-30Bibliographically approved
Haseeb, S., Javed, S., Mokayed, H., Martin-del-Campo, S., Sandin, F., Liwicki, M. & Delsing, J. (2026). Local Cloud-based Collaborative Learning vs Other IIoT Decentralized AI Solutions: A Systematic Literature Review. Journal of Network and Systems Management, Article ID 49.
Open this publication in new window or tab >>Local Cloud-based Collaborative Learning vs Other IIoT Decentralized AI Solutions: A Systematic Literature Review
Show others...
2026 (English)In: Journal of Network and Systems Management, ISSN 1064-7570, E-ISSN 1573-7705, article id 49Article, review/survey (Refereed) Published
Abstract [en]

The increasing complexity and dynamic nature of Industrial Internet of Things (IIoT) demand scalable, adaptive, intuitive, and real-time automation frameworks. This paper presents a systematic literature review (SLR) of edge- and cloud-based collaborative learning frameworks for predictive maintenance and smart manufacturing tasks. In this SLR, we highlight the under-utilization of distributed computational architectures that provide complete automation support (design and run-time), flexibility, scalability, and inter- & intra-cloud service exchange while adhering to security management and integrity principles for solving IIoT tasks using modern artificial intelligence (AI) models at the edge/cloud. Recently, many IIoT applications have been designed using AI models that require robust, low-latency, and data-secure frameworks. This demand drives a trend toward distributed computational architectures in which data storage and processing are partially or fully decentralized. Common paradigms addressing this resource distribution include edge computing, federated learning, and private or hybrid clouds. We analyze 50 recent studies against IoT characteristics, AI performance metrics, and network/system management requirements. Our findings reveal underutilization of distributed architectures that support automation, interoperability, and security. While most solutions rely on centralized or hybrid clouds, fewer than 5% adopt federated or transfer learning, and over 60% remain dependent on supervised models. We also introduce a comparative perspective on network and security management, showing that local/private cloud implementations can reduce control-plane overhead and synchronization latency, though gaps persist in dynamic bandwidth allocation and zero-trust adoption. Finally, we benchmark our previously proposed local cloud-based collaborative learning (CCL) model against state-of-the-art solutions, highlighting its strengths in automation and interoperability, as well as limitations in adaptive computation and intelligent offloading. This review identifies the research gaps and opportunities for integrating collaborative AI, secure automation, and hybrid architectures to meet Industry 5.0 objectives of resilience, sustainability, and human-centricity. 

Place, publisher, year, edition, pages
Springer Nature, 2026
Keywords
Systematic literature review (SLR), Industrial internet of things (IIoT), Edge AI, Cloud AI, Federated learning, Predictive maintenance, Smart manufacturing, Cloud-based architectures, Local cloud, Unsupervised learning, Collaborative learning
National Category
Computer Sciences Computer Systems Communication Systems
Research subject
Machine Learning; Cyber-Physical Systems
Identifiers
urn:nbn:se:ltu:diva-111751 (URN)10.1007/s10922-025-10029-y (DOI)001689347500001 ()2-s2.0-105030028851 (Scopus ID)
Projects
Arrowhead flexible Production Value Network (fPVN)
Funder
European Commission, 101111977
Note

Funder: AI-REDGIO5.0 (101092069);

Full text license: CC BY;

This article has previously appeared as a manuscript in a thesis.

Available from: 2025-02-25 Created: 2025-02-25 Last updated: 2026-06-30Bibliographically approved
Haseeb, S., Minhas, A. M., Liwicki, F. S., Gardelli, V., Liwicki, M., Stamouli, P., . . . Mokayed, H. (2026). Q2A (Question-To-Answer): A Mixed-Methods Evaluation Pipeline For Ai-Supported Educational Tools: Evidence From The Ai4Edu Project. In: Proceedings of EDULEARN26 Conference: . Paper presented at 18th International Conference on Education and New Learning Technologies, Palma, Spain, 29 June-1 July, 2026.. IATED, Article ID 1347.
Open this publication in new window or tab >>Q2A (Question-To-Answer): A Mixed-Methods Evaluation Pipeline For Ai-Supported Educational Tools: Evidence From The Ai4Edu Project
Show others...
2026 (English)In: Proceedings of EDULEARN26 Conference, IATED , 2026, article id 1347Conference paper, Published paper (Refereed) [Artistic work]
Abstract [en]

Large Language Model (LLM)-based conversational assistants are increasingly entering classrooms, yet systematic evidence remains limited regarding the educational themes they effectively support and how users perceive and engage with them. The rapid deployment of artificial intelligence (AI) tools in education has introduced a critical methodological challenge: how can heterogeneous feedback data, such as interviews, Likert-scale questionnaires, open-ended survey responses, and system usage logs, be systematically integrated into a coherent, research-aligned evaluation process? While collecting feedback is common practice, structured methodologies for synthesizing diverse qualitative and quantitative evidence into collective conclusions remain underdeveloped. To address this gap, we introduce Q2A (Question-to-Answer), a three-stage mixed-methods evaluation pipeline developed within the AI4EDU project to assess two GenAI applications designed to support independent learning and instructional delivery: Study Buddy (student-facing) and Teacher Mate (teacher-facing) in upper secondary education. The Q2A pipeline consists of: (1) a human-centered, LLM-assisted mapping method aligning survey items with predefined research questions to ensure construct validity; (2) integrated qualitative and quantitative analyses, including thematic clustering of Likert-scale constructs, sentiment analysis (TextBlob, VADER, and BERT) of open-ended responses, and behavioral analytics derived from system logs; and (3) an aggregation stage that synthesizes findings across multiple analyses to produce unified, research-aligned conclusions. The Q2A pipeline was implemented across four European countries (Sweden, Greece, Cyprus, and Ireland), involving approximately 30 teachers and over 300 students. In the first stage, LLMs (GPT and Gemini) were used to generate mapping between specific questions from the survey instruments and the research questions designed by the project evaluators. These mappings were then reviewed and structured by human pedagogical experts, who also appended additional data sources to the mapping that should be considered for answering the research question(s). For example, evaluating the impact of Study Buddy on student motivation required aggregating survey responses, teacher feedback, usage logs, and academic performance indicators. The final aggregation stage revealed moderate to high positive effects on student motivation (3.7/5 in Sweden and Cyprus; 4.2/5 in Greece and Ireland). Importantly, results demonstrated that perceived educational value emerged from the convergence of survey data and behavioral indicators, highlighting how the mixed-methods aggregation of multiple feedback data sources provides greater interpretive clarity than isolated metrics alone. The Q2A pipeline offers a replicable and scalable framework for integrating multi-source educational data into systematic evaluation processes, providing practical guidance for researchers and policymakers assessing AI-supported learning environments.

Place, publisher, year, edition, pages
IATED, 2026
Keywords
Large Language Models (LLMs), Artificial Intelligence in Education (AIED), Conversational AI, GenAI, Sentiment Analysis, Thematic Analysis, Lexicon/Rule-Based Sentiment Analysis, Transformer-Based Sentiment Analysis, Educational Technology Assessment, AI4EDU
National Category
Computer Sciences Natural Language Processing
Research subject
Machine Learning; Education
Identifiers
urn:nbn:se:ltu:diva-119224 (URN)10.21125/edulearn.2026.1347 (DOI)
Conference
18th International Conference on Education and New Learning Technologies, Palma, Spain, 29 June-1 July, 2026.
Projects
GenAI-EduMap
Note

Funder: European Union;

ISBN for host publication: 978-84-09-88444-5

Available from: 2026-08-11 Created: 2026-08-11 Last updated: 2026-09-02Bibliographically approved
Shehzadi, T., Ifza, I., Liwicki, M., Stricker, D. & Afzal, M. Z. (2026). Semi-Supervised Object Detection: A Survey on Progress from CNN to Transformer. Sensors, 26, Article ID 310.
Open this publication in new window or tab >>Semi-Supervised Object Detection: A Survey on Progress from CNN to Transformer
Show others...
2026 (English)In: Sensors, E-ISSN 1424-8220, Vol. 26, article id 310Article, review/survey (Refereed) Published
Abstract [en]

The impressive advancements in semi-supervised learning have driven researchers to explore its potential in object detection tasks within the field of computer vision. Semi-Supervised Object Detection (SSOD) leverages a combination of a small labeled dataset and a larger, unlabeled dataset. This approach effectively reduces the dependence on large labeled datasets, which are often expensive and time-consuming to obtain. Initially, SSOD models encountered challenges in effectively leveraging unlabeled data and managing noise in generated pseudo-labels for unlabeled data. However, numerous recent advancements have addressed these issues, resulting in substantial improvements in SSOD performance. This paper presents a comprehensive review of 28 cutting-edge developments in SSOD methodologies, from Convolutional Neural Networks (CNNs) to Transformers. We delve into the core components of semi-supervised learning and its integration into object detection frameworks, covering data augmentation techniques, pseudo-labeling strategies, consistency regularization, and adversarial training methods. Furthermore, we conduct a comparative analysis of various SSOD models, evaluating their performance and architectural differences. We aim to ignite further research interest in overcoming existing challenges and exploring new directions in semi-supervised learning for object detection.

Place, publisher, year, edition, pages
MDPI, 2026
Keywords
transformer, object detection, DETR, computer vision, deep neural networks
National Category
Computer graphics and computer vision
Research subject
Machine Learning
Identifiers
urn:nbn:se:ltu:diva-116087 (URN)10.3390/s26010310 (DOI)001657672700001 ()41516746 (PubMedID)2-s2.0-105027136243 (Scopus ID)
Note

Full text license: CC BY 4.0;

Funder: AIRISE (101092312)

Available from: 2026-01-21 Created: 2026-01-21 Last updated: 2026-06-30Bibliographically approved
Dorigo, T., Brown, G. D., Casonato, C., Cerda, A., Ciarrochi, J., Lio, M. D., . . . Yazdanpanah, N. (2025). Artificial Intelligence in Science and Society: the Vision of USERN. IEEE Access, 13, 15993-16054
Open this publication in new window or tab >>Artificial Intelligence in Science and Society: the Vision of USERN
Show others...
2025 (English)In: IEEE Access, E-ISSN 2169-3536, Vol. 13, p. 15993-16054Article, review/survey (Refereed) Published
Abstract [en]

The recent rise in relevance and diffusion of Artificial Intelligence (AI)-based systems and the increasing number and power of applications of AI methods invites a profound reflection on the impact of these innovative systems on scientific research and society at large. The Universal Scientific Education and Research Network (USERN), an organization that promotes initiatives to support interdisciplinary science and education across borders and actively works to improve science policy, collects here the vision of its Advisory Board members, together with a selection of AI experts, to summarize how we see developments in this exciting technology impacting science and society in the foreseeable future. In this review, we first attempt to establish clear definitions of intelligence and consciousness, then provide an overviewof AI’s state of the art and its applications. A discussion of the implications, opportunities, and liabilities of the diffusion of AI for research in a few representative fields of science follows this. Finally, we address the potential risks of AI to modern society, suggest strategies for mitigating those risks, and present our conclusions and recommendations.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers Inc., 2025
National Category
Computer Sciences
Research subject
Machine Learning; Pervasive Mobile Computing
Identifiers
urn:nbn:se:ltu:diva-111443 (URN)10.1109/ACCESS.2025.3529357 (DOI)001410357500037 ()2-s2.0-85215252598 (Scopus ID)
Note

Validerad;2025;Nivå 2;2025-03-12 (u2);

Full text license: CC BY;

For funding information, see: 10.1109/ACCESS.2025.3529357

Available from: 2025-01-28 Created: 2025-01-28 Last updated: 2025-10-21Bibliographically approved
Chhipa, P. C., Vashishtha, G., Anantha sai Settur, J., Saini, R., Shah, M. & Liwicki, M. (2025). ASTrA: Adversarial Self-supervised Training with Adaptive-Attacks. In: ASTrA: Adversarial Self-supervised Training with Adaptive-Attacks: . Paper presented at 13th International Conference on Learning Representations (ICLR 2025), Singapore, Republic of Singapore, April 24-28, 2025 (pp. 100735-100757).
Open this publication in new window or tab >>ASTrA: Adversarial Self-supervised Training with Adaptive-Attacks
Show others...
2025 (English)In: ASTrA: Adversarial Self-supervised Training with Adaptive-Attacks, 2025, p. 100735-100757Conference paper, Poster (with or without abstract) (Refereed)
Abstract [en]

Existing self-supervised adversarial training (self-AT) methods rely on hand-crafted adversarial attack strategies for PGD attacks, which fail to adapt to the evolving learning dynamics of the model and do not account for instance-specific characteristics of images. This results in sub-optimal adversarial robustness and limits the alignment between clean and adversarial data distributions. To address this, we propose ASTrA (Adversarial Self-supervised Training with Adaptive-Attacks), a novel framework introducing a learnable, self-supervised attack strategy network that autonomously discovers optimal attack parameters through exploration-exploitation in a single training episode. ASTrA leverages a reward mechanism based on contrastive loss, optimized with REINFORCE, enabling adaptive attack strategies without labeled data or additional hyperparameters. We further introduce a mixed contrastive objective to align the distribution of clean and adversarial examples in representation space. ASTrA achieves state-of-the-art results on CIFAR10, CIFAR100, and STL10 while integrating seamlessly as a plug-and-play module for other self-AT methods. ASTrA shows scalability to larger datasets, demonstrates strong semi-supervised performance, and is resilient to robust overfitting, backed by explainability analysis on optimal attack strategies. Project page for source code and other details at https://prakashchhipa.github.io/projects/ASTrA.

National Category
Computer Vision and Learning Systems Artificial Intelligence
Research subject
Machine Learning
Identifiers
urn:nbn:se:ltu:diva-111564 (URN)2-s2.0-105010233560 (Scopus ID)
Conference
13th International Conference on Learning Representations (ICLR 2025), Singapore, Republic of Singapore, April 24-28, 2025
Note

Full text license: CC BY

Available from: 2025-02-07 Created: 2025-02-07 Last updated: 2026-02-12Bibliographically approved
Liwicki, F. S., Saini, R., Das Chakladar, D., Rakesh, S., Gupta, V., Liwicki, M. & Eriksson, J. (2025). Dataset: Synchronous EEG and fMRI dataset on inner speech. Luleå University of Technology
Open this publication in new window or tab >>Dataset: Synchronous EEG and fMRI dataset on inner speech
Show others...
2025 (English)Other (Other academic)
Abstract [en]

This dataset contains simultaneous EEG-fMRI recordings for inner speech experiments. Data were collected using a 3T MRI scanner and 64-channel BrainProducts EEG system. The EEG data have undergone preprocessing, including pulse artifact removal, using the BrainVision Analyzer software. No further data transformations have been applied to ensure the dataset remains BIDS-compliant as "raw".

Place, publisher, year, pages
Luleå University of Technology, 2025
National Category
Computer Sciences Artificial Intelligence
Research subject
Machine Learning
Identifiers
urn:nbn:se:ltu:diva-115637 (URN)10.18112/openneuro.ds006033.v1.0.1 (DOI)
Funder
Luleå University of Technology, LTU-154-2023, LTU-4908-2022The Kempe Foundations, JCSMK23-0102
Note

Full text license: CC0;

Repository: OpenNeuro;

Related item(s): DOI 10.1016/j.dib.2025.112258 (Data article); 

Available from: 2025-12-03 Created: 2025-12-03 Last updated: 2025-12-08Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-4029-6574

Search in DiVA

Show all publications