Salim Sazzed
Published in: HT ’22: Proceedings of the 33rd ACM Conference on Hypertext and Social Media · DOI: 10.1145/3511095.3536358
License: © 2022 Copyright held by the owner/author(s).
ACM source attribution. Complete text transcribed from the ACM Hypertext 2022 proceedings HTML, verified against the companion proceedings PDF, under ACM’s confirmed authorization. The terminal Visual-Meta appendix is publisher metadata and excluded from the scholarly article body.
Revealing the Demographic Attributes of the Authors from the Abstracts of Scientific Articles
Salim Sazzed Old Dominion University · Norfolk, VA, USA · ssazz001@odu.edu
ABSTRACT
This study presents multiple strategies to automatically reveal undisclosed demographic attributes of the authors in the double-blind submissions. From a limited amount of textual content of around 100-200 words excerpted from an abstract, this study aims to reveal the following pieces of information, i) the English language nativeness of the primary author, ii) the country of origin of the primary author, and iii) the gender of the primary author. We introduce an annotated dataset of over 5600 articles labeled with the native language, country of origin, and gender information of the primary authors. We employ classical machine learning (CML) algorithms with statistical n-gram features and transformer-based fine-tuned language models to determine various demographic attributes. We observe that transformer-based models yield slightly better performances for all three tasks. The transformer-based models achieve macro F1 scores close to 75% for identifying the English language nativeness of the primary authors. To determine the country of the non-native English authors, the fine-tuned transformer-based models obtain F1 scores of around 60% (10-class classification). For the gender prediction task, we attain F1 scores of 0.65 by the transformer-based models. The experimental results demonstrate that the fine-tuned language models and CML classifiers are capable of disclosing various author attributes with an acceptable level of accuracy that can undermine the blindness of the double-blind submission.
CCS CONCEPTS
• Social and professional topics → User characteristics; • Information systems → World Wide Web.
KEYWORDS
demography prediction, native language identification, scientific publication, gender prediction, peer review, article abstract, authorship attribution
ACM Reference Format: Salim Sazzed. 2022. Revealing the Demographic Attributes of the Authors from the Abstracts of Scientific Articles. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT ’22), June 28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3511095.3536358
1 INTRODUCTION
To avoid preconceptions in the peer-review process, often, the authors are instructed to conceal identifiable information such as names and university or country affiliations. Although the reviewers are expected to maintain the highest integrity and refrain from exercising any preferential or adversarial attitudes, due to human nature, implicit bias may occur. A reasonable solution to avoid bias in the review process is to keep the author's details hidden from the reviewers, known as the double-blind review process, as opposed to the single-blind review process where only the reviewer's information is kept confidential [ 13 ].
Stylometry studies linguistic styles to discover inferential properties of documents, such as various author attributes. The stylometric analysis helps reveal demographic information crucial for decision-making in various domains. Thus, there has been a growing interest in applying stylometry for various tasks such as determining deceptive online reviews [ 17 ], gender identification [ 2 , 11 ], native language identification [ 12 , 20 , 21 ]. A number of studies employed stylometric analysis for identifying authors in scientific writings from a given or probable author list [ 4 , 8 , 10 ]. These studies either utilized i) the full text of the articles or ii) the list of references mentioned in the articles or both i) and ii) to identify authors from the given list.
Unlike the previous studies [ 1 , 5 , 22 ], here, we focus on a more generalizable goal of determining demographic attributes of authors that may cause implicit bias in the peer-review process. We leverage only the text available in an article's abstract, usually fewer than 200 words. Using this limited textual content, we aim to uncover various demographic attributes of the primary author in the double-blind submission. The primary author of an article is the author who has the supervisory role in the research design, writing, and publication process.
In particular, we aim to answer the following research questions from the short text excerpted from the abstract-
RQ1: Can we predict whether a paper was written by a native or non-native English author (i.e., English language nativeness of the primary author)?
RQ2: Is it possible to infer the country of origin of the native and non-native English authors?
RQ3: Is the gender information of the primary author deducible from the abstract?
We develop a novel annotated corpus consisting of over 5600 research articles labeled with various demographic attributes of the primary authors. The corpus represents publications of natural language processing (NLP) and the computational linguistic (CL) domains and collected from the ACL Anthology, the largest digital archive of NLP and computational linguistic (CL) articles. We manually annotate all the articles with the attributes required for the following three identification tasks: i) the English language nativeness of the primary author, ii) the native country of the primary author, and iii) the gender of the primary author.
Besides, we analyze various synthetic attributes of the abstracts, such as the mean sentence length and frequency of verb/adjectives/articles representing authors of different groups that are important from a sociolinguistic standpoint. To automatically perform all three tasks, we employ classical ML (CML) classifiers with word n-gram based tf-idf features. Besides, we fine-tune transformer-based pre-trained language models BERT (Bidirectional Encoder Representations from Transformers) and RoBERTa (Robustly optimized BERT) for the demography prediction tasks. We observe that transformer-based fine-tuned models yield slightly better efficacy for all classification tasks.
1.1 Contributions, Significance and Novelty of the Proposed Study
1.1.1 Contributions. We introduce a large dataset of over 5600 conference and workshop papers annotated with the English language nativeness, the country of origin, and the gender of the primary authors and make the dataset publicly available 1 . We leverage CML classifiers and fine-tuned transformer-based language models and demonstrate that it is possible to infer various demographic attributes of authors with an acceptable level of accuracy utilizing a short amount of text of fewer than 200 words.
1.1.2 Significance: The outcomes of this study are critical to comprehending the usefulness of the double-blind submission process. Many top conferences (e.g., ICWSM) provide the option to choose manuscripts to review from a pool of anonymized abstracts. The reviewers may choose manuscripts representing authors of particular demography if demographic information is inferable from the abstract, which undermines the effectiveness of the double-blind peer review system. To the best of our knowledge, this is the first study that demonstrates the possibility of the scenario that ML classifiers can disclose demographic information from the abstract.
1.1.3 Novelty: As far as we are aware, none of the existing work considered the following aspects and goals while determining the demography of authors of scientific articles-
Use of fine-tuned language models: Fine-tuned language models have demonstrated remarkable performances for various classification tasks; thus, in this study, we aim to find the effectiveness of the fine-tuned language models for inferring author demographic information from the abstract of the scientific articles.
Country of native English speakers: The task of identifying the country of English native authors has not been explored in earlier works. Here, we demonstrate that it is possible to differentiate between native English authors of English-speaking countries, such as the USA, UK, and AU, to some extent.
Abstract only information: Most importantly, here, we use limited information retrieved from an article's abstract to infer author demography. While existing studies tried to identify the author(s) of a scientific article from a given list of authors and published articles, here, we focus on a more generalizable goal of determining author demographic attributes, which may add implicit bias in the peer-review process.
2 DATASETS
2.1 Data Collection
All the data utilized in this study are collected from the ACL Anthology [ 18 ]. We download the BibTeX data representing the collection of articles available in the ACL anthology. The BibTeX data contains various types of information such as the paper title, venue, list of authors, years of publications, and abstracts. Since we require only the titles and abstracts of the papers and names of the authors, a python script is written to extract only the required information.
2.2 Data Annotation and Statistics
The data annotation task consists of labeling articles with various pieces of information, such as the native languages, countries, and genders of the primary authors. Usually, among the list of authors, the first and last author play the most influential roles [ 3 , 19 , 23 ]. The last author is usually the corresponding author who supervises the entire project and is deemed to have the most influential role in the final writing [ 23 ].
Although, an exception from this typical scenario is possible when the supervisor is both the first and corresponding author. To address both kinds of scenarios, we count the number of publications by both first and last authors in the ACL Anthology. Since ACL Anthology archives articles from all the top CL and NLP publication venues (e.g., ACL, EMNLP, NAACL), it is possible to get comprehensive information regarding the scholarly reputations and experiences of the authors. We select the author with a higher scholarly reputation (i.e.,#publications,#years of publications) as the primary author. The annotations of the articles regarding language, country, and gender-specific information are performed considering the primary author. As we want to maintain high accurateness in the labeling process, all the author's information is manually verified from the author's websites. In case of the missing websites of the authors or if we are not convinced about the author's demography from the available data, we do not include the abstract in the dataset.
Table 1: The country-specific distributions of articles representing non-native and native English authors.
Native | Non-Native | ||||
|---|---|---|---|---|---|
Country | #Pub. | Percentage(%) | Country | #Pub. | Percentage(%) |
China (CN) | 616 | 20.99% | Australia (AU) | 151 | 5.64% |
Japan (JP) | 490 | 16.70% | USA (US) | 1780 | 66.49% |
Germany (DK) | 689 | 23.48% | UK (UK) | 643 | 24.01% |
Italy (IT) | 122 | 4.15% | Canada (CA) | 103 | 3.84% |
Spain (ES) | 155 | 5.28% | |||
Israel (IL) | 144 | 4.90% | |||
India (IN) | 245 | 8.35% | |||
France (FR) | 248 | 8.45% | |||
Netherland (NL) | 118 | 4.02% | |||
Hongkong (HK) | 107 | 3.64%, | |||
Total | 2934 | 100% | Total | 2677 | 100% |
2.2.1 Native vs. Non-Native English Speakers: We perform manual labeling following a strict set of criteria. In addition to the name of the author, we consider two additional measures, i) the country of current affiliation and ii) the undergrad institution. Note that the author's current affiliation may not resemble the author's country of origin. For example, many international researchers work in the USA as faculty. Thus, considering only the country or university affiliation will not suffice to determine the English language nativeness of the primary author. The country-specific annotation for a primary author is performed considering all the three criteria; An author is included in the dataset only if all three criteria refer to the same country. The authors who belong to any of the predominantly English-speaking countries are labeled as native English speakers.
2.2.2 Country of Non-native and Native English Authors. We further label the country of native and non-native English authors following similar criteria of the author's country of affiliation, name, and undergrad institution. Authors from the following non-native English countries are considered in this study: Germany, China, Japan, India, Spain, Israel, Netherland, France, Italy, and Hongkong. These countries represent the most number of publications by non-native English authors in the ACL anthology, thus, used in this study. The authors from the following four English-native countries: USA, UK, Australia (AU), and Canada (CA), are considered. Few other English native countries, such as New Zealand and Ireland, are not considered in this study as they have comparatively fewer authors in the ACL Anthology.
2.2.3 Gender of Author. To determine the gender-specific information of the primary author of an article, we look for any available photo on the author's website. Although the author's name may provide gender-specific information, it may not be accurate always. If we can not find any photo of the primary author or we can not retrieve the gender-specific information of the author, we do not include that author and corresponding articles in the dataset.
The final dataset contains over 5600 papers representing primary authors from 14 countries (Table 1 ). The linguistic attributes of various groups are shown in Table 2 . Note that in this study, we only consider binary-level gender.
Table 2: Statistics of various syntactic attributes in abstracts of two groups of authors (English native and English non-native) and (Male and Female), respectively.
Attribute | Native | Non-native |
|---|---|---|
#Publications | 2677 | 2934 |
Mean sentence length (#words) | 22.72 | 21.71 |
Mean abstract length (#words) | 132.33 | 130.87 |
#Articles per abstract (mean) | 9.39 | 9.95 |
#Adjectives per abstract (mean) | 16.55 | 15.52 |
#Verbs per abstract (mean) | 17.62 | 16.85 |
Attribute | Male | Female |
#Publications | 4382 | 1257 |
Mean sentence length (#words) | 22.16 | 22.32 |
Mean abstract length (#words) | 131.33 | 132.28 |
#Articles per abstract (mean) | 9.75 | 9.38 |
#Adjectives per abstract (mean) | 16.01 | 16.21 |
#Verbs per abstract (mean) | 17.29 | 17.44 |
3 CLASSICAL ML CLASSIFIERS AND PRE-TRAINED LANGUAGE MODELS
We employ four classical ML (CML) classifiers: Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), and K-Nearest Neighbor (k-NN) for predicting various demographic attributes of the primary authors. We extract the word unigrams and bigrams from the abstracts and compute their tf-idf scores which are then fed to the classifiers.
Table 3: Performance of various classifiers for predicting the English language nativenes of the authors.
Type | Classifier | Precision | Recall | F1 | Accuracy |
|---|---|---|---|---|---|
LR | 0.735 | 0.735 | 0.735 | 0.73 | |
CML | SVM | 0.735 | 0.735 | 0.735 | 0.73 |
RF | 0.689 | 0.689 | 0.689 | 0.68 | |
KNN | 0.67 | 0.66 | 0.66 | 0.66 | |
Transformer | BERT | 0.745 | 0.743 | 0.744 | 0.74 |
ROBERTA | 0.750 | 0.735 | 0.742 | 0.73 |
Table 4: The performances of various classifiers for predicting the country of non-native English authors (10-class classification task). ’-’ indicates undefined F1 score.
Method | CN | JP | DK | IT | ES | IL | IN | FR | NL. | HK | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | |
LR | 0.73/0.86 | 0.64/0.66 | 0.63/0.75 | 0.55/0.48 | 0.43/0.32 | 0.47/0.34 | 0.70/0.67 | 0.64/0.48 | 0.30/0.19 | 0.22/0.13 | 0.58/0.63 |
SVM | 0.73/0.8 | 0.59/0.48 | 0.54/0.9 | 0.26/0.15 | 0.09/0.05 | 0.17/0.1 | 0.58/0.43 | 0.62/0.46 | -/0.01 | -/0.01 | -/0.55 |
RF | 0.65/0.8 | 0.53/0.39 | 0.52/0.87 | -/0.07 | -/0.0 | -/0.02 | 0.36/0.23 | 0.62/0.46 | -/0.01 | -/0.0 | -/0.50 |
KN | 0.6/0.82 | 0.54/0.56 | 0.57/0.56 | 0.45/0.39 | 0.38/0.29 | 0.28/0.18 | 0.52/0.46 | 0.71/0.59 | 0.32/0.23 | 0.27/0.19 | 0.50/0.54 |
BERT | 0.80/0.74 | 0.69/0.84 | 0.66/0.72 | 0.63/0.5 | 0.46/0.4 | 0.38/0.29 | 0.77/0.75 | 0.86/0.79 | 0.42/0.45 | 0.29/0.20 | 0.61/0.68 |
RoBERTa | 0.75/0.75 | 0.71/0.84 | 0.65/0.62 | 0.50/0.33 | 0.37/0.33 | 0.56/0.50 | 0.67/0.75 | 0.85/0.83 | 0.56/0.64 | 0.4/0.3 | 0.62/0.67 |
Table 5: Performances of classifiers for predicting the country of native English speakers (4-class).
Classifier | US | UK | AU | CA | Total |
|---|---|---|---|---|---|
F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | F1/Acc. | |
LR | 0.82/0.91 | 0.5/0.42 | 0.24/0.15 | 0.36/0.22 | 0.54/0.72 |
SVM | 0.82/0.94 | 0.5/0.4 | -/0.04 | 0.11/0.06 | -/0.73 |
RF | 0.8/1.0 | 0.06/0.03 | -/0.0 | -/0.0% | -/0.67 |
KNN | 0.8/0.88 | 0.47/0.41 | 0.16/0.1 | 0.42/0.28 | 0.51/0.70 |
BERT | 0.82/0.83 | 0.58/0.59 | 0.27/0.27 | 0.46/0.30 | 0.57/0.72 |
RoBERTa | 0.83/0.82 | 0.58/0.60 | 0.26/0.28 | 0.45/0.30 | 0.56/0.72 |
Table 6: Performances of various classifiers for the gender prediction task.
Classifier | Female | Male | Total | |
|---|---|---|---|---|
F1/Acc. | F1/Acc. | F1/Acc. | ||
LR | 0.43 / 0.41 | 0.85 / 0.86 | 0.64 / 0.76 | |
SVM | 0.35 / 0.26 | 0.87 / 0.93 | 0.63 / 0.78 | |
RF | 0.35 / 0.26 | 0.87 / 0.93 | 0.63 / 0.78 | |
KNN | 0.35 / 0.26 | 0.87 / 0.94 | 0.63 / 0.78 | |
BERT | 0.48 / 0.56 | 0.82 / 0.78 | 0.65 /0.73 | |
RoBERTa | 0.46 / 0.53 | 0.83 / 0.80 | 0.64 /0.74 |
We employ two variants of transformer-based language models, BERT [ 6 ] and RoBERTa [ 16 ]. We fine-tune the pre-trained models for classifying abstracts into two or multiple classes based on the task considered. Since all the tasks in this study are classification problems, we utilize the classification module of pre-trained models. The hugging face library [ 25 ] is used for fine-tuning all the pre-trained models. Only the last layers of the pre-trained models are fine-tuned for the classification tasks. A mini-batch size of 16 and a learning rate of 0.00003 are used. During the training, 20% samples are utilized as a validation set. The Adam optimizer is used for the optimization, and the loss function is set to sparse-categorical-cross-entropy. The training process runs for 4 epochs, and an early stopping criterion is employed.
4 RESULTS AND DISCUSSION
We use 10-fold cross-validation to assess the performances of various approaches. As seen by Table 3 , the transformer-based BERT and ROBERTA yield the best performances by attaining macro F1 scores of around 0.75 for predicting the English language nativeness of the primary authors. Among all the tasks, ML classifiers yield the best performances for this English language nativeness identification task which suggests that some distinctions exist in abstracts written by native and non-native English authors.
Table 4 displays the performances of various classifiers for predicting the native language of non-native English authors. This dataset is highly class-imbalanced, which affects the performances of the classifiers across different classes. For example, the most dominant non-native English author group (i.e., DK) holds 23.5% of the samples in the dataset, while the smallest group HK has less than 3.5% samples in the dataset. We observe that the results of the classifiers are heavily biased towards the dominant class. Among the various CML classifiers, the best accuracy is obtained by LR.
To address the class imbalance issue, we set the class-weight to balanced in scikit-learn implementations of all the CML classifiers. Although it improves the performance for the minor class prediction, it still suffers from the class-imbalanced issue. For example, most of the classifiers (if not all) show poor performance for predicting authors from HK, NL, and IL classes, as these classes have less than 150 samples in the dataset. We investigate the performance of several oversampling and under-sampling techniques such as SMOTE from the imbalanced-learn package [ 14 ]. However, we obtain similar or worse results with more computational time than using the ’balanced’ class settings of the scikit-learn library. Thus, the reported results for the CML classifiers in Table 4 are based on the class-balanced settings. For the fine-tuned transformer-based models, we find that over-sampling the minor classes yields the best macro F1 scores.
Table 5 provides the performances of ML classifiers for determining the country of native English authors. Similar to the non-native dataset, this dataset is also class-imbalanced. All the CML classifiers perform poorly for predicting two minor classes, AU and CA. Unsurprisingly, ML classifiers show the best performance for predicting US-based authors, who have over 70% presence in the dataset.
For the gender prediction task (Table 6 ), we obtain the best macro F1 score of 0.65 by the transformer-based BERT model. The relatively low proportion of the articles from female authors in the dataset affects the performance of the female author identification. The mediocre performances of ML classifiers for this binary-level author gender prediction task suggest that the gender prediction task is very challenging.
Although some of the earlier work applied synthetic features such as frequency of function words [ 24 ], sentence length [ 7 ], and PoS [ 9 ] for the authorship attribution or groups of authors prediction tasks in the social media text, we find due to the essence of our data these synthetic features do not provide enough clues for our prediction tasks. Since abstracts of scientific articles represent the instances of formal writing of scholarly people, signals that exist in social media text or descriptive essays (e.g., frequency of sentiment words, usage of common adjectives and verbs, and grammatical errors) may not be present here. Besides, other style words such as punctuation, emojis, and slang words are also not available in the abstract of scientific articles.
An interesting observation is that ML classifiers yield worse performance for predicting the country of native English authors than the country of non-native English authors, although the first task is a 4-classification problem, while the latter task is a 10-class classification problem. The best ML classifier achieves F1 scores of 0.57 and 0.62, respectively, for these two tasks. The influences of the native language of non-native English authors are reflected in their English writing, which assists ML classifiers to distinguish non-native English authors better. In contrast, such distinction is less prevalent in native English author groups (USA, UK, AUS) as they have a common root. Even though there exist differences between USA, UK, or AU English, those are less apparent in the short abstract of scientific articles.
Note that the presence of outliers (e.g., abstract modified by a professional English translation service) or non-standard practice (e.g., author-ordering) in the dataset is not impossible. However, any such kind of presence has minimal impact on the results due to the low frequency of such cases. Data annotation for any kind of task always contain uncertainty to some extent. Note that this study does not claim that the non-native English authors are less fluent in written English than the native-English authors; instead, it reveals that there exist some distinguishing signals between the languages of the native and non-native authors even within a brief amount of text available in published abstracts. Besides, the binary-level gender classification does not reflect any of the author's perceptions regarding gender; instead, it follows the gender classification approach of existing works such as [ 2 , 15 , 26 ].
5 SUMMARY AND CONCLUSIONS
In this study, we investigate whether we can reveal the demographic attributes of the primary authors from the limited text available in the abstracts of double-blind submissions. In particular, we aim to identify the following author attributes: the English language nativeness, country, and gender, by incorporating labeled data into supervised ML classifiers. To accomplish this goal, we first create an annotated dataset of over 5600 articles marked with the native language, country, and gender-specific information of the primary authors. We apply various CML classifiers with word n-gram features for the prediction tasks. Besides, we fine-tune the transformer-based pre-trained language models to predict the demographic attributes of the authors. Our results and finding suggested that given a reasonable amount of labeled data, ML classifiers can unveil various author information of the double-blind submissions. Although the purpose of the double-blind review process is to maintain complete anonymity to avoid any review bias, our findings show that this goal may not be fully attainable.
REFERENCES
[1] Douglas Bagnall. 2015. Author identification using multi-headed recurrent neural networks. arXiv preprint arXiv: 1506.04891 (2015).
[2] David Bamman, Jacob Eisenstein, and Tyler Schnoebelen. 2014. Gender identity and lexical variation in social media. Journal of Sociolinguistics 18, 2 (2014), 135–160.
[3] Surajit Bhattacharya et al. 2010. Authorship issue explained. Indian J Plast Surg 43, 2 (2010), 233–4.
[4] Cornelia Caragea, Ana Uban, and Liviu P Dinu. 2019. The myth of double-blind review revisited: ACL vs. EMNLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 2317–2327.
[5] Stephen J Ceci and Douglas Peters. 1984. How blind is blind review? American Psychologist 39, 12 (1984), 1491.
[6] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv: 1810.04805 (2018).
[7] Gili Goldin, Ella Rabinovich, and Shuly Wintner. 2018. Native language identification with user generated content. In Proceedings of the 2018 conference on empirical methods in natural language processing . 3591–3601.
[8] Shawndra Hill and Foster Provost. 2003. The myth of the double-blind review? Author identification using only citations. Acm Sigkdd Explorations Newsletter 5, 2 (2003), 179–184.
[9] Graeme Hirst and Ol'ga Feiguina. 2007. Bigrams of syntactic labels for authorship discrimination of short texts. Literary and Linguistic Computing 22, 4 (2007), 405–417.
[10] Julian Hitschler, Esther Van Den Berg, and Ines Rehbein. 2017. Authorship attribution with convolutional neural networks and POS-eliding. In Proceedings of the Workshop on Stylistic Variation . 53–58.
[11] Moshe Koppel, Shlomo Argamon, and Anat Rachel Shimoni. 2002. Automatically categorizing written texts by author gender. Literary and linguistic computing 17, 4 (2002), 401–412.
[12] Moshe Koppel, Jonathan Schler, and Kfir Zigdon. 2005. Determining an author's native language by mining a text for errors. In Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining . 624–628.
[13] Carole J Lee, Cassidy R Sugimoto, Guo Zhang, and Blaise Cronin. 2013. Bias in peer review. Journal of the American Society for Information Science and Technology 64, 1 (2013), 2–17.
[14] Guillaume Lemaître, Fernando Nogueira, and Christos K. Aridas. 2017. Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. Journal of Machine Learning Research 18, 17 (2017), 1–5. http://jmlr.org/papers/v18/16-365.html
[15] Wen Li and Markus Dickinson. 2017. Gender prediction for Chinese social media data. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017 . 438–445.
[16] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv: 1907.11692 (2019).
[17] Myle Ott, Yejin Choi, Claire Cardie, and Jeffrey T Hancock. 2011. Finding deceptive opinion spam by any stretch of the imagination. arXiv preprint arXiv: 1107.4557 (2011).
[18] Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara. 2013. The ACL anthology network corpus. Language Resources and Evaluation 47, 4 (2013), 919–944.
[19] Ruchita Sarawgi, Kailash Gajulapalli, and Yejin Choi. 2011. Gender attribution: tracing stylometric evidence beyond topic and genre. In Proceedings of the fifteenth conference on computational natural language learning . 78–86.
[20] Salim Sazzed. 2021. A Hybrid Approach of Opinion Mining and Comparative Linguistic Analysis of Restaurant Reviews. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021) . 1281–1288.
[21] Salim Sazzed. 2022. Influence of Language Proficiency on the Readability of Review Text and Transformer-based Models for Determining Language Proficiency. (2022).
[22] Andrew Tomkins, Min Zhang, and William D Heavlin. 2017. Reviewer bias in single-versus double-blind peer review. Proceedings of the National Academy of Sciences 114, 48(2017), 12708–12713.
[23] Teja Tscharntke, Michael E Hochberg, Tatyana A Rand, Vincent H Resh, and Jochen Krauss. 2007. Author sequence and credit for contributions in multiauthored publications. PLoS biology 5, 1 (2007), e18.
[24] Vered Volansky, Noam Ordan, and Shuly Wintner. 2015. On the features of translationese. Digital Scholarship in the Humanities 30, 1 (2015), 98–118.
[25] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, Online, 38–45. https://www.aclweb.org/anthology/2020.emnlp-demos.6
[26] Chuhan Wu, Fangzhao Wu, Tao Qi, Junxin Liu, Yongfeng Huang, and Xing Xie. 2019. Neural gender prediction in microblogging with emotion-aware user representation. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management . 2401–2404.
FOOTNOTE
FOOTNOTE 1 https://github.com/sazzadcsedu/DemographyScientificArticle.git Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). HT '22, June 28–July 01, 2022, Barcelona, Spain © 2022 Copyright held by the owner/author(s). ACM ISBN 978-1-4503-9233-4/22/06. DOI: https://doi.org/10.1145/3511095.3536358
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime