Integrating Manual and Automatic Annotation for the Creation of Discourse Network Data Sets

My bibliography Save this article

Integrating Manual and Automatic Annotation for the Creation of Discourse Network Data Sets

Author

Listed:

Sebastian Haunss
(Research Center on Inequality and Social Policy, University of Bremen, Germany)
Jonas Kuhn
(Institute for Natural Language Processing, University of Stuttgart, Germany)
Sebastian Padó
(Institute for Natural Language Processing, University of Stuttgart, Germany)
Andre Blessing
(Institute for Natural Language Processing, University of Stuttgart, Germany)
Nico Blokker
(Research Center on Inequality and Social Policy, University of Bremen, Germany)
Erenay Dayanik
(Institute for Natural Language Processing, University of Stuttgart, Germany)
Gabriella Lapesa
(Institute for Natural Language Processing, University of Stuttgart, Germany)

Registered:

Sebastian Haunss

Abstract

This article investigates the integration of machine learning in the political claim annotation workflow with the goal to partially automate the annotation and analysis of large text corpora. It introduces the MARDY annotation environment and presents results from an experiment in which the annotation quality of annotators with and without machine learning based annotation support is compared. The design and setting aim to measure and evaluate: a) annotation speed; b) annotation quality; and c) applicability to the use case of discourse network generation. While the results indicate only slight increases in terms of annotation speed, the authors find a moderate boost in annotation quality. Additionally, with the help of manual annotation of the actors and filtering out of the false positives, the machine learning based annotation suggestions allow the authors to fully recover the core network of the discourse as extracted from the articles annotated during the experiment. This is due to the redundancy which is naturally present in the annotated texts. Thus, assuming a research focus not on the complete network but the network core, an AI-based annotation can provide reliable information about discourse networks with much less human intervention than compared to the traditional manual approach.

Suggested Citation

Sebastian Haunss & Jonas Kuhn & Sebastian Padó & Andre Blessing & Nico Blokker & Erenay Dayanik & Gabriella Lapesa, 2020. "Integrating Manual and Automatic Annotation for the Creation of Discourse Network Data Sets," Politics and Governance, Cogitatio Press, vol. 8(2), pages 326-339.

Handle: RePEc:cog:poango:v8:y:2020:i:2:p:326-339
DOI: 10.17645/pag.v8i2.2591

Download full text from publisher

References listed on IDEAS

Melanie Nagel & Keiichi Satoh, 2019. "Protesting iconic megaprojects. A discourse network analysis of the evolution of the conflict over Stuttgart 21," Urban Studies, Urban Studies Journal Limited, vol. 56(8), pages 1681-1700, June.
Laver, Michael & Benoit, Kenneth & Garry, John, 2003. "Extracting Policy Positions from Political Texts Using Words as Data," American Political Science Review, Cambridge University Press, vol. 97(2), pages 311-331, May.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

Philip Leifeld, 2020. "Policy Debates and Discourse Network Analysis: A Research Agenda," Politics and Governance, Cogitatio Press, vol. 8(2), pages 180-183.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Sebastian Haunss & Jonas Kuhn & Sebastian Padó & Andre Blessing & Nico Blokker & Erenay Dayanik & Gabriella Lapesa, 2020. "Integrating Manual and Automatic Annotation for the Creation of Discourse Network Data Sets," Politics and Governance, Cogitatio Press, vol. 8(2), pages 326-339.
Armèn Hakhverdian, 2009. "Capturing Government Policy on the Left–Right Scale: Evidence from the United Kingdom, 1956–2006," Political Studies, Political Studies Association, vol. 57(4), pages 720-745, December.
Simon Hug & Tobias Schulz, 2007. "Referendums in the EU’s constitution building process," The Review of International Organizations, Springer, vol. 2(2), pages 177-218, June.
Pongsak Luangaram & Yuthana Sethapramote, 2016. "Central Bank Communication and Monetary Policy Effectiveness: Evidence from Thailand," PIER Discussion Papers 20, Puey Ungphakorn Institute for Economic Research.
Rybinski, Krzysztof, 2020. "The forecasting power of the multi-language narrative of sell-side research: A machine learning evaluation," Finance Research Letters, Elsevier, vol. 34(C).
Hamza Bennani, 2012. "National influences inside the ECB: an assessment from central bankers' statements," Working Papers hal-00992646, HAL.
Kenneth Benoit & Michael Laver & Christine Arnold & Paul Pennings & Madeleine O. Hosli, 2005. "Measuring National Delegate Positions at the Convention on the Future of Europe Using Computerized Word Scoring," European Union Politics, , vol. 6(3), pages 291-313, September.
Weiss, Max & Zoorob, Michael, 2021. "Political frames of public health crises: Discussing the opioid epidemic in the US Congress," Social Science & Medicine, Elsevier, vol. 281(C).
Yang, Chao & Huang, Cui, 2022. "Quantitative mapping of the evolution of AI policy distribution, targets and focuses over three decades in China," Technological Forecasting and Social Change, Elsevier, vol. 174(C).
Matthew Gentzkow & Jesse M. Shapiro & Matt Taddy, 2019. "Measuring Group Differences in High‐Dimensional Choices: Method and Application to Congressional Speech," Econometrica, Econometric Society, vol. 87(4), pages 1307-1340, July.
- Matthew Gentzkow & Jesse M. Shapiro & Matt Taddy, 2016. "Measuring Group Differences in High-Dimensional Choices: Method and Application to Congressional Speech," NBER Working Papers 22423, National Bureau of Economic Research, Inc.
Shengli Dai & Weimin Zhang & Jiamin Zong & Yingying Wang & Ge Wang, 2021. "How Effective Is the Green Development Policy of China’s Yangtze River Economic Belt? A Quantitative Evaluation Based on the PMC-Index Model," IJERPH, MDPI, vol. 18(14), pages 1-17, July.
Ralf Meinhardt & Sebastian Junge & Martin Weiss, 2018. "The organizational environment with its measures, antecedents, and consequences: a review and research agenda," Management Review Quarterly, Springer, vol. 68(2), pages 195-235, April.
Torsten J. Selck, 2005. "Improving the Explanatory Power of Bargaining Models," Journal of Theoretical Politics, , vol. 17(3), pages 371-375, July.
Cory Koedel & Jiaxi Li & Matthew G. Springer & Li Tan, 2018. "Teacher Performance Ratings and Professional Improvement," Working Papers 1808, Department of Economics, University of Missouri.
Kayla N. Jordan & James W. Pennebaker & Chase Ehrig, 2018. "The 2016 U.S. Presidential Candidates and How People Tweeted About Them," SAGE Open, , vol. 8(3), pages 21582440187, July.
Michael Evans & Wayne McIntosh & Jimmy Lin & Cynthia Cates, 2007. "Recounting the Courts? Applying Automated Content Analysis to Enhance Empirical Legal Research," Journal of Empirical Legal Studies, John Wiley & Sons, vol. 4(4), pages 1007-1039, December.
Hayo, Bernd & Henseler, Kai & Rapp, Marc Steffen, 2019. "Estimating the monetary policy interest-rate-to-performance sensitivity of the European banking sector at the zero lower bound," Finance Research Letters, Elsevier, vol. 31(C).
- Bernd Hayo & Kai Henseler & Marc Steffen Rapp, 2019. "Estimating the monetary policy interest-rate-to-performance sensitivity of the European banking sector at the zero lower bound," MAGKS Papers on Economics 201902, Philipps-Universität Marburg, Faculty of Business Administration and Economics, Department of Economics (Volkswirtschaftliche Abteilung).
Huang, Cui & Yang, Chao & Su, Jun, 2021. "Identifying core policy instruments based on structural holes: A case study of China’s nuclear energy policy," Journal of Informetrics, Elsevier, vol. 15(2).
Hanna Bäck & Marc Debus & Wolfgang C. Müller, 2016. "Intra-party diversity and ministerial selection in coalition governments," Public Choice, Springer, vol. 166(3), pages 355-378, March.
Sarel, Roee & Demirtas, Melanie, 2021. "Delegation in a multi-tier court system: Are remands in the U.S. federal courts driven by moral hazard?," European Journal of Political Economy, Elsevier, vol. 68(C).
- Sarel, Roee & Demirtas, Melanie, 2019. "Delegation in a multi-tier court system: are remands in the U.S. federal courts driven by moral hazard?," ILE Working Paper Series 28, University of Hamburg, Institute of Law and Economics.

More about this item

Keywords

annotation; automation; discourse networks; machine learning; migration discourse;
All these keywords.

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:cog:poango:v8:y:2020:i:2:p:326-339. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: António Vieira or IT Department (email available below). General contact details of provider: https://www.cogitatiopress.com .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Integrating Manual and Automatic Annotation for the Creation of Discourse Network Data Sets

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data