Special Issue of the Language Resources and Evaluation Journal

Language Resources and Evaluation for Healthcare

Recent advances in Natural Language Processing (NLP), Speech Technologies, and Large Language Models (LLMs) have significantly expanded the range of language-based applications in healthcare. Language technologies are increasingly being used to support the analysis of electronic health records (EHRs) (Da Rocha et al. 2022), the creation of specialized datasets and corpora (e Oliveira et al. 2022; Santos, Oliveira, and Paraboni 2024), the development of domain-specific language models (Gumiel et al. 2019; Nunes et al. 2024), and downstream applications such as named entity recognition (Andrade, Ruas, and Couto 2021; Schneider et al. 2020), information extraction, clinical decision support, and conversational agents (Pires, Caseli, and Neris 2023; de Souza et al. 2022). More recently, foundation models and multimodal approaches have further broadened the potential impact of language technologies in areas such as mental health, patient communication, accessibility, and public health.

Despite this rapid progress, significant challenges remain. The development of robust and trustworthy health-related language technologies depends on the availability of high-quality language resources, including corpora, benchmarks, annotation schemes, lexical resources, and evaluation frameworks. Additional challenges involve multilinguality, data scarcity, privacy preservation, de-identification, reproducibility, fairness, and the reliable evaluation of increasingly complex language models. Addressing these challenges requires not only advances in algorithms and models, but also the creation, documentation, sharing, and systematic evaluation of language resources.

This Special Issue aims to bring together researchers and practitioners working at the intersection of language technologies and healthcare. We welcome contributions addressing the creation, extension, evaluation, and application of language resources for health-related NLP, as well as studies on benchmarking, reproducibility, responsible AI, and the evaluation of emerging foundation models in healthcare settings.

The Special Issue is motivated by the success of the First Workshop on Language Technologies for Health (Lang4Health), held in conjunction with PROPOR 2026. While extended and substantially revised versions of papers presented at the workshop are welcome, the Special Issue is open to new submissions from the broader international community.

Although contributions involving any language are encouraged, we particularly welcome work on Portuguese, Galician, and other under-resourced languages. The initiative builds upon the growing body of research in health-related language technologies developed by the Portuguese-speaking community (Schneider et al. 2020; e Oliveira et al. 2022; Gumiel et al. 2019; Da Rocha et al. 2022; Nunes et al. 2024; Andrade, Ruas, and Couto 2021; Souza et al. 2020; Oliveira and Paraboni 2024; Consoli et al. 2022), while fostering broader international collaboration and resource sharing across languages and healthcare domains.

Target Audience

The Special Issue will solicit papers both from among the participants of the First Workshop on Language Technologies for Health (Lang4Health) and from the wider community via an open call for contributions, targeting in particular contributors from among related workshops such as The Third Workshop on Patient-oriented Language Processing (CL4Health) and The 8th Clinical Natural Language Processing Workshop, and also among the communities in relevant mailing-lists (Corpora, SigLex, ELRA, ACL Portal). We particularly welcome contributions addressing the creation, annotation, documentation, distribution, and evaluation of language resources for health-related applications, as well as benchmarking studies, reproducibility analyses, and evaluation methodologies for language and speech technologies in healthcare. While the initiative originated from Lang4Health, the Special Issue is open to submissions from the broader international research community, with particular interest in studies involving Portuguese, Galician, and other under-resourced languages.

Call for Paper

Special Issue on Language Resources and Evaluation for Healthcare

Recent advances in Natural Language Processing (NLP), Speech Technologies, and Large Language Models (LLMs) have transformed the development of language-based applications in healthcare. From the analysis of electronic health records and patient-generated content to conversational agents, decision-support systems, and digital health applications, language technologies are increasingly being deployed in real-world healthcare settings.

The success of these technologies, however, depends critically on the availability of high-quality language resources and rigorous evaluation methodologies. Healthcare applications pose unique challenges, including specialized terminology, multilinguality, data scarcity, privacy constraints, annotation complexity, reproducibility, and the need for trustworthy and explainable systems. As language technologies continue to evolve, there is a growing need for resources, benchmarks, evaluation frameworks, and methodological studies that support reliable and responsible research and development in this domain.

This Special Issue aims to bring together researchers and practitioners working at the intersection of language resources, evaluation, and healthcare. The initiative builds on discussions from the First Workshop on Language Technologies for Health (Lang4Health) and welcomes contributions from the broader international community.

We invite original contributions addressing the creation, annotation, documentation, dissemination, evaluation, and application of language resources for healthcare-related language and speech technologies. Extended versions of papers presented at Lang4Health are welcome, provided they contain substantial new material. New submissions are equally encouraged.

Topics of interest include, but are not limited to:

  • Data in health domain
    • Dataset/Corpus construction and availability;
    • Dataset/Corpus annotation;
    • Anonymization and de-identification;
    • Augment and synthetic data generation.
  • Language technologies in health domain
    • Interaction and conversational agents (e.g. chatbots);
    • Information extraction and information retrieval;
    • Named Entity Recognition;
    • Summarization;
    • Question Answering;
    • Personalization;
    • Speech processing;
    • NLP-supported diagnosis;
    • Accessibility and simplification of information;
    • Language support for digital phenotyping.

Submissions involving any language are welcome. We particularly encourage contributions addressing Portuguese, Galician, and other under-resourced languages.

Manuscripts must present original and unpublished research and will undergo the standard peer-review process of the Language Resources and Evaluation journal.

Submission information

Submissions must present original and unpublished research that is not under consideration elsewhere. We welcome substantial contributions that advance the development, evaluation, documentation, dissemination, or application of language resources and evaluation methodologies for healthcare-related language and speech technologies.

Extended versions of papers presented at the First Workshop on Language Technologies for Health (Lang4Health) are welcome, provided that they contain substantial new material beyond the workshop publication. New submissions from the broader research community are equally encouraged.

Contributions may include, but are not limited to, new language resources, resource extensions, annotation methodologies, evaluation frameworks, benchmarking studies, reproducibility analyses, and innovative applications of language resources in healthcare settings.

Manuscripts should be prepared and submitted according to the guidelines of the Language Resources and Evaluation journal. Detailed submission instructions will be provided upon the launch of the Special Issue.

Review process

All submissions will undergo the standard -blind peer-review process of the Language Resources and Evaluation journal. Each manuscript will be evaluated by at least two expert reviewers selected by the Guest Editors based on their expertise and the absence of conflicts of interest. Reviewers will assess submissions with respect to scientific quality, originality, methodological soundness, relevance to the Special Issue, and potential impact on the field. To support a rigorous and timely review process, authors submitting to the Special Issue may be invited to participate in the reviewer pool and review other submissions within their area of expertise. Reviewer assignments will be managed by the Guest Editors to ensure appropriate expertise and to avoid conflicts of interest. Revised manuscripts may undergo additional rounds of review when necessary.

Important dates

  • Deadline for paper submission: October 2nd, 2026
  • Notification for authors: December 15, 2026

Guest Editors

  • Prof. Aline Villavicencio
    University of Exeter, UK

  • Dr. Rodrigo Wilkens
    University of Exeter, UK  

  • Prof. Helena Caseli
    Federal University of São Carlos, Brazil

  • Dr. Vânia Neris
    Federal University of São Carlos, Brazil

Contact information

Email

References

Andrade, Vitor DT, Pedro Ruas, and Francisco M Couto. 2021. “Named Entity Recognition and Linking: A Portuguese and Spanish Oncological Parallel Corpus.” bioRxiv, 2021–09.
Consoli, Bernardo, Henrique D. P. dos Santos, Ana Helena D. P. S. Ulbrich, Renata Vieira, and Rafael H. Bordini. 2022. BRATECA (Brazilian Tertiary Care Dataset): A Clinical Information Dataset for the Portuguese Language.” In Proceedings of the Thirteenth Language Resources and Evaluation Conference, edited by Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, et al., 5609–16. Marseille, France: European Language Resources Association. https://aclanthology.org/2022.lrec-1.602/.
Da Rocha, Naila Camila, Abner Macola Pacheco Barbosa, Yaron Oliveira Schnr, Juliana Machado-Rugolo, Luis Gustavo Modelli de Andrade, José Eduardo Corrente, and Liciana Vaz de Arruda Silveira. 2022. “Natural Language Processing to Extract Information from Portuguese-Language Medical Records.” Data 8 (1): 11.
de Souza, Paula Maia, Isabella da Costa Pires, Vivian Genaro Motti, Helena Medeiros Caseli, Jair Barbosa Neto, Larissa C Martini, and Vânia Paula de Almeida Neris. 2022. “Design Recommendations for Chatbots to Support People with Depression.” In Proceedings of the 21st Brazilian Symposium on Human Factors in Computing Systems. IHC ’22. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3554364.3559119.
e Oliveira, Lucas Emanuel Silva, Ana Carolina Peters, Adalniza Moura Pucca Da Silva, Caroline Pilatti Gebeluca, Yohan Bonescki Gumiel, Lilian Mie Mukai Cintho, Deborah Ribeiro Carvalho, Sadid Al Hasan, and Claudia Maria Cabral Moro. 2022. “SemClinBr-a Multi-Institutional and Multi-Specialty Semantically Annotated Corpus for Portuguese Clinical NLP Tasks.” Journal of Biomedical Semantics 13 (1): 13.
Gumiel, Yohan Bonescki, Arnon Bruno Ventrilho dos Santos, Lilian Mie Mukai Cintho, Deborah Ribeiro Carvalho, Sadid A Hasan, Claudia Maria Cabral Moro, et al. 2019. “Learning Portuguese Clinical Word Embeddings: A Multi-Specialty and Multi-Institutional Corpus of Clinical Narratives Supporting a Downstream Biomedical Task.” In MEDINFO 2019: Health and Wellbeing e-Networks for All, 123–27. IOS Press.
Nunes, Miguel, João Boné, João C Ferreira, Pedro Chaves, and Luis B Elvas. 2024. “MediAlbertina: An European Portuguese Medical Language Model.” Computers in Biology and Medicine 182: 109233.
Oliveira, Rafael, and Ivandré Paraboni. 2024. “A Bag-of-Users Approach to Mental Health Prediction from Social Media Data.” In Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, edited by Pablo Gamallo, Daniela Claro, António Teixeira, Livy Real, Marcos Garcia, Hugo Gonçalo Oliveira, and Raquel Amaro, 509–14. Santiago de Compostela, Galicia/Spain: Association for Computational Lingustics. https://aclanthology.org/2024.propor-1.52/.
Pires, Isabella, Helena Caseli, and Vânia Neris. 2023. “Design de Um Chatbot Para o Diálogo Com Universitários Com Possível Perfil Depressivo.” In Anais Estendidos Do XXIII Simpósio Brasileiro de Computação Aplicada à Saúde, 7–12. Porto Alegre, RS, Brasil: SBC. https://doi.org/10.5753/sbcas_estendido.2023.229543.
Santos, Wesley Ramos dos, Rafael Lage de Oliveira, and Ivandré Paraboni. 2024. SetembroBR: a social media corpus for depression and anxiety disorder prediction.” Language Resources and Evaluation 58 (1): 273–300. https://doi.org/10.1007/s10579-022-09633-0.
Schneider, Elisa Terumi Rubel, João Vitor Andrioli de Souza, Julien Knafou, Lucas Emanuel Silva e Oliveira, Jenny Copara, Yohan Bonescki Gumiel, Lucas Ferro Antunes de Oliveira, Emerson Cabrera Paraiso, Douglas Teodoro, and Cláudia Maria Cabral Moro Barra. 2020. BioBERTpt - a Portuguese Neural Language Model for Clinical Named Entity Recognition.” In Proceedings of the 3rd Clinical Natural Language Processing Workshop, edited by Anna Rumshisky, Kirk Roberts, Steven Bethard, and Tristan Naumann, 65–72. Online: Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.clinicalnlp-1.7.
Souza, João Vitor Andrioli de, Elisa Terumi Rubel Schneider, Josilaine Oliveira Cezar, Lucas Emanuel Silva, Yohan Bonescki Gumiel, Emerson Cabrera Paraiso, Douglas Teodoro, Claudia Maria Cabral Moro Barra, et al. 2020. “A Multilabel Approach to Portuguese Clinical Named Entity Recognition.” Journal of Health Informatics 12.