Fine-Grained Emotion Classification from Mobile App Reviews: A

Fine-Grained Emotion Classification from Mobile App Reviews: A 2026 Empirical Study with Large Language Models


arXiv:2610.03802 | Submitted 1 October 2026


Authors: Quim Motger, Carlota Catot, Marc Oriol


Subjects: Computation and Language (cs.CL); Information Retrieval (cs.IR); Software Engineering (cs.SE)


Cite as: arXiv:2610.03802 [cs.CL]




Abstract


Context. Fine-grained emotion classification of mobile app reviews enables requirements engineering (RE) activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied β€” a gap that has become increasingly pressing as app store feedback volumes continue to grow and as LLM-based RE pipelines mature in 2026.


Objectives. Building on a previously published annotation framework and human-labelled ground truth adapted from Plutchik's taxonomy, this paper investigates how large language models (LLMs) can be leveraged for automatic multi-label emotion classification under severe class imbalance.


Methods. The study compares:


  • Encoder-only fine-tuning under multi-label and binary-ensemble formulations
  • Decoder-only zero- and few-shot prompting across open-source and proprietary models
  • A catalogue of imbalance mitigation strategies, including loss reweighting, resampling, and generative data augmentation

The synthetic-review generator and prompting strategy are selected via an intrinsic augmentation-utility ranking.


Results. Fine-tuned encoders trail the best decoder-only few-shot prompting (macro-F1 0.642) by a wide margin at baseline (multi-label: 0.387; binary ensemble: 0.450). Pairing the best multi-label encoder with generative data augmentation and positive-weighted loss closes most of this gap (+0.204), operating at up to three orders of magnitude lower inference latency than the decoders. The largest gains appear on the rarest emotions, moving from undetected to improvements of up to +0.501 F1.


Conclusion. LLMs make fine-grained, multi-label emotion classification of app reviews feasible for requirements engineering pipelines, albeit with modest macro-F1. The best formulation and mitigation strategy prove backbone- and formulation-dependent. The authors release the experimental pipeline, synthetic corpora, and fine-tuned checkpoints for replication and reuse.




Key Contributions


  • A systematic comparison of encoder-only and decoder-only LLM formulations for multi-label emotion classification in app reviews.
  • A catalogue of imbalance mitigation strategies, empirically tested under severe class imbalance.
  • An intrinsic augmentation-utility ranking for selecting synthetic-review generators and prompting strategies.
  • Open artifacts: experimental pipeline, synthetic corpora, and fine-tuned checkpoints released for replication and reuse.



Why This Matters in 2026


As mobile applications continue to dominate digital service delivery, app store reviews have become one of the richest sources of user sentiment and requirement-relevant feedback. Moving beyond binary polarity toward Plutchik-style fine-grained emotion detection gives product and RE teams a more actionable signal for prioritising issues and analysing feature-level responses. This work shows that LLM-based formulations β€” particularly few-shot decoding and augmentation-enhanced encoder fine-tuning β€” can operationalise such signals today, while flagging the trade-offs between accuracy, latency, and backbone choice that practitioners must weigh.




Availability


The experimental pipeline, synthetic corpora, and fine-tuned checkpoints are released by the authors to support replication and downstream reuse.


DOI: https://doi.org/10.48550/arXiv.2610.03802

via ArXiv CL+LG

Related