On-Device Named-Entity Recognition: A Deployability Study of

On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence


Author: Vinay Kumar Chaganti

arXiv: arXiv:2610.00007 [cs.CL]

Submitted: 9 July 2026

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

ACM Class: I.2.7

Comments: 7 pages, 5 figures, 12 tables. Code and per-span records reproduce all reported numbers offline


Abstract


Named-entity recognition (NER) is increasingly wanted on-device—no API, low latency, data kept local. The practitioner's question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. This work answers these questions jointly.


The study places nine systems across three paradigms and 13M to 8B parameters:


  • a classical tagger (spaCy),
  • bidirectional-encoder specialists (GLiNER, 166M to 460M),
  • generative LLMs run locally (Qwen3-0.6B/1.7B/4B-Instruct, DeepSeek-R1-1.5B/8B).

These are evaluated on three datasets of differing character. The paper reports accuracy plus two axes often omitted in the literature: latency and output validity.


Because the corpus (RSS-News) had no gold labels, the authors built silver gold from a cross-family LLM judge panel, then measured its fidelity against benchmark gold and a full human re-validation of the corpus. Strict F1 reached 0.95—an upper bound, since the human gold was silver-seeded. Gold provenance flips the paradigm ranking: moving from LLM-authored silver to human gold raises every encoder and lowers every generative model.


On accuracy alone, a 4B instruct LLM is competitive—it leads on clean newswire—so the encoder's case is one of deployability. It matches or slightly trails at one-ninth to one-twenty-fourth the size, at millisecond-to-second latency, with zero malformed output. Meanwhile, the smallest generative models emit up to 27% invalid output on long inputs, a failure fixed by scale, not output budget.


The paper then characterizes GLiNER's per-span confidence: it ranks correctness well (AUROC 0.76 to 0.86) but is overconfident (ECE 0.24 to 0.47, halved by temperature scaling). Thresholding gives a small honest out-of-sample F1 gain. An all-local small-to-large cascade gives a modest, corpus-dependent gain over cost-matched random routing. Confidence tracks correctness but not novelty.


Every number recomputes offline from per-span records.


Key Findings


  • Deployability over leaderboard accuracy: Encoder models offer strong practical advantages in size, latency, and output validity despite generative LLMs being competitive on clean newswire.
  • Evaluation without human annotation: A cross-family LLM judge panel can produce silver gold with high fidelity (strict F1 0.95) when validated against benchmark and human re-validation.
  • Gold provenance matters: The ranking of paradigms changes depending on whether evaluation uses LLM-authored silver or human gold, with encoders improving and generative models declining under human gold.
  • Output validity: The smallest generative models produce up to 27% invalid output on long inputs; this is mitigated by scale rather than output budget.
  • Confidence calibration: GLiNER's per-span confidence is discriminative (AUROC 0.76–0.86) but overconfident (ECE 0.24–0.47); temperature scaling halves ECE.
  • Cascade routing: A small-to-large local cascade yields modest, corpus-dependent gains over cost-matched random routing.
  • Reproducibility: All results are reproducible offline from per-span records.

via ArXiv CL+LG

Related