ESPERANTO: Evaluating Synthesized Phrases to Enhance Robustness in AI Detection for Text Origination
While large language models (LLMs) exhibit significant utility across various domains, they simultaneously are susceptible to exploitation for unethical purposes, including academic misconduct and dissemination of misinformation. Consequently, AI-generated text detection systems have emerged as a countermeasure. However, these detection mechanisms demonstrate vulnerability to evasion techniques and lack robustness against textual manipulations. This paper introduces back-translation as a novel technique for evading detection, underscoring the need to enhance the robustness of current detection systems. The proposed method involves translating AI-generated text through multiple languages before back-translating to English. We present a model that combines these back-translated texts to produce a manipulated version of the original AI-generated text. Our findings demonstrate that the manip
doi
10.1145/3720553.3746665
isbn
979-8-4007-1534-1
name
ESPERANTO: Evaluating Synthesized Phrases to Enhance Robustness in AI Detection for Text Origination
source
supplied BITS XML and PDF
acm_url
https://dl.acm.org/doi/10.1145/3720553.3746665
authors
Navid Ayoobi, Lily Knab, Wen Cheng, David Pantoja, Hamidreza Alikhani, Sylvain Flamant, Jin Kim, Arjun Mukherjee
doi_url
https://doi.org/10.1145/3720553.3746665
license
© 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
summary
While large language models (LLMs) exhibit significant utility across various domains, they simultaneously are susceptible to exploitation for unethical purposes, including academic misconduct and dissemination of misinformation. Consequently, AI-generated text detection systems have emerged as a countermeasure. However, these detection mechanisms demonstrate vulnerability to evasion techniques and lack robustness against textual manipulations. This paper introduces back-translation as a novel technique for evading detection, underscoring the need to enhance the robustness of current detection systems. The proposed method involves translating AI-generated text through multiple languages before back-translating to English. We present a model that combines these back-translated texts to produce a manipulated version of the original AI-generated text. Our findings demonstrate that the manip
keywords
(empty)
arxiv_url
https://arxiv.org/abs/2409.14285
published
2025-09-15
conference
HT '25: 36th ACM Conference on Hypertext and Social Media, Chicago, IL, USA, September 15–19, 2025
open_access
false
acm_html_url
https://dl.acm.org/doi/full/10.1145/3720553.3746665
ccs_concepts
(empty)
displayAuthor
Navid Ayoobi, Lily Knab, Wen Cheng, David Pantoja, Hamidreza Alikhani, Sylvain Flamant, Jin Kim, Arjun Mukherjee
proceedings_url
https://dl.acm.org/doi/proceedings/10.1145/3720553
displayPublishTime
2025-09-15
acm_reference_format
(empty)