ConfliBERT-Arabic: A Pre-trained Arabic Language Model for Politics, Conflicts and Violence

Sultan Alsarra, Luay Abdeljaber, Wooseong Yang, Niamat Zawad, Latifur Khan, Patrick T. Brandt, Javier Osorio, Vito J. D'Orazio

Research output: Chapter in Book/Report/Conference proceedingConference contribution

1 Scopus citations

Abstract

This study investigates the use of Natural Language Processing (NLP) methods to analyze politics, conflicts and violence in the Middle East using domain-specific pre-trained language models. We introduce Arabic text and present ConfliBERT-Arabic, a pre-trained language models that can efficiently analyze political, conflict and violence-related texts. Our technique hones a pre-trained model using a corpus of Arabic texts about regional politics and conflicts. Performance of our models is compared to baseline BERT models. Our findings show that the performance of NLP models for Middle Eastern politics and conflict analysis are enhanced by the use of domain-specific pre-trained local language models. This study offers political and conflict analysts, including policymakers, scholars, and practitioners new approaches and tools for deciphering the intricate dynamics of local politics and conflicts directly in Arabic.

Original languageEnglish (US)
Title of host publicationInternational Conference Recent Advances in Natural Language Processing, RANLP 2023
Subtitle of host publicationLarge Language Models for Natural Language Processing - Proceedings
EditorsGalia Angelova, Maria Kunilovskaya, Ruslan Mitkov
PublisherIncoma Ltd
Pages98-108
Number of pages11
ISBN (Electronic)9789544520922
DOIs
StatePublished - 2023
Event2023 International Conference Recent Advances in Natural Language Processing: Large Language Models for Natural Language Processing, RANLP 2023 - Varna, Bulgaria
Duration: Sep 4 2023Sep 6 2023

Publication series

NameInternational Conference Recent Advances in Natural Language Processing, RANLP
ISSN (Print)1313-8502

Conference

Conference2023 International Conference Recent Advances in Natural Language Processing: Large Language Models for Natural Language Processing, RANLP 2023
Country/TerritoryBulgaria
CityVarna
Period9/4/239/6/23

ASJC Scopus subject areas

  • Software
  • Computer Science Applications
  • Artificial Intelligence
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'ConfliBERT-Arabic: A Pre-trained Arabic Language Model for Politics, Conflicts and Violence'. Together they form a unique fingerprint.

Cite this