Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Instruction Makes a Difference
Luleå University of Technology, Department of Computer Science, Electrical and Space Engineering, Embedded Internet Systems Lab.ORCID iD: 0000-0002-5582-2031
Luleå University of Technology, Department of Computer Science, Electrical and Space Engineering, Embedded Internet Systems Lab.ORCID iD: 0009-0005-9310-1701
Luleå University of Technology, Department of Computer Science, Electrical and Space Engineering, Embedded Internet Systems Lab.ORCID iD: 0000-0003-1343-1742
Luleå University of Technology, Department of Computer Science, Electrical and Space Engineering, Embedded Internet Systems Lab.ORCID iD: 0000-0003-2039-3844
2024 (English)In: Document Analysis Systems: 16th IAPR International Workshop, DAS 2024, Athens, Greece, August 30–31, 2024, Proceedings / [ed] Giorgos Sfikas; George Retsinas, Springer Science and Business Media Deutschland GmbH , 2024, p. 71-88Conference paper, Published paper (Refereed)
Abstract [en]

We introduce the Instruction Document Visual Question Answering (iDocVQA) dataset and the Large Language Document (LLaDoc) model, for training Language-Vision (LV) models for document analysis and predictions on document images, respectively. Usually, deep neural networks for the DocVQA task are trained on datasets lacking instructions. We show that using instruction-following datasets improves performance. We compare performance across document-related datasets using the recent state-of-the-art (SotA) Large Language and Vision Assistant (LLaVA)1.5 as the base model. We also evaluate the performance of the derived models for object hallucination using the Polling-based Object Probing Evaluation (POPE) dataset. The results show that instruction-tuning performance ranges from 11x to 32x of zero-shot performance and from 0.1% to 4.2% over non-instruction (traditional task) finetuning. Despite the gains, these still fall short of human performance (94.36%), implying there’s much room for improvement.

Place, publisher, year, edition, pages
Springer Science and Business Media Deutschland GmbH , 2024. p. 71-88
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 14994
Keywords [en]
DocVQA, instruction-tuning, LLM, LMM
National Category
Computer Sciences Computer graphics and computer vision
Research subject
Machine Learning
Identifiers
URN: urn:nbn:se:ltu:diva-110169DOI: 10.1007/978-3-031-70442-0_5ISI: 001334866300005Scopus ID: 2-s2.0-85204640516OAI: oai:DiVA.org:ltu-110169DiVA, id: diva2:1903573
Conference
16th IAPR International Workshop on Document Analysis Systems (DAS 2024), Athens, Greece, August 30-31, 2024
Funder
Knut and Alice Wallenberg Foundation
Note

Funder: Wallenberg AI, AutonomousSystems and Software Program (WASP)

ISBN for host publication: 978-3-031-70441-3, 978-3-031-70442-0

Available from: 2024-10-04 Created: 2024-10-04 Last updated: 2025-10-21Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Adewumi, TosinHabib, NudratAlkhaled, LamaBarney, Elisa

Search in DiVA

By author/editor
Adewumi, TosinHabib, NudratAlkhaled, LamaBarney, Elisa
By organisation
Embedded Internet Systems Lab
Computer SciencesComputer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 346 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf