On the Limits of LLM Adaptability

Impact of Model-Internalized Priors on Annotation Task Performance

Etienne Casanova*, Rafal Kocielnik*, R. Michael Alvarez
California Institute of Technology
*Equal contribution · Spotlight, Oral Presentation · ICML 2026
Study overview: given an input message and a task definition, an LLM produces a label. RQ1 asks whether task familiarity predicts performance; RQ2 asks whether prompting can rescue zero-shot errors; RQ3 asks how misaligned definitions affect predictions and confidence.

Figure 1. Given a message and a user-provided task definition, the LLM predicts a label. We study three facets of how the model's internalized concept of the task interacts with that instruction: familiarity (RQ1), steerability under an aligned definition (RQ2), and susceptibility to a misaligned one (RQ3).

+0.41 partial r

Definition-Specific Familiarity predicts accuracy after controlling for dataset; text memorization does not (r = −0.15 to −0.19).

34.8% rescue rate

of zero-shot errors are corrected by definitions, examples, or DSPy optimization. Nearly two-thirds resist correction entirely.

17% vs 5% definition vs model

Definition wording swings accuracy by up to 17 points across models; model choice alone swings it by only 5.

Abstract

Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions relates to performance, (2) whether additional information in prompts can correct zero-shot errors (“decision stickiness”), and (3) model susceptibility to misaligned task definitions. We introduce Definition-Specific Familiarity (DSF), which measures alignment between a model's elicited concept and the target definition. Across nine LLMs and five diverse toxicity datasets (spanning social media, gaming, news, and forums), DSF predicts model annotation performance after controlling for dataset identity (partial r = +0.41), and this association remains positive across all six prompting conditions tested. In contrast, three common text-memorization metrics show no positive association. We show that prompting has limited corrective power: only 34.8% of zero-shot errors are corrected by additional instructions or examples, with high-confidence errors especially persistent. Misaligned definitions systematically shift predictions without reducing reported confidence, making confidence unreliable for detecting definition–policy mismatch. Together, these findings establish definition alignment as a practical model-selection criterion and show that better prompting alone cannot substitute for validating model–policy fit.

RQ1

Definition fit predicts performance, memorization doesn't

Does alignment between a model's internalized concept and a task definition predict which model performs better, beyond text memorization and dataset differences?

Two annotators handed the same guideline can still disagree, because each brings a different prior conception of the concept. We treat LLMs the same way. To compute Definition-Specific Familiarity (DSF), we ask a model to explain its own understanding of the target concept (“In your own words, what makes content toxic?”) and measure the semantic similarity between that explanation and the dataset's official definition, averaged across six independent sentence encoders (MiniLM, MPNet, BGE-large, E5-large, Instructor-large, OpenAI text-embedding-3-small).

This is different from asking whether the model has memorized the dataset's text, which we measure separately with ROUGE-L, BERTScore, and embedding similarity on prefix-continuation pairs. The two signals diverge sharply:

Chart of partial correlations with accuracy, controlling for dataset, across six prompting conditions. Consensus DSF is positive in every condition, from about +0.24 to +0.62. ROUGE-L, BERTScore, and embedding similarity memorization metrics cluster near zero or slightly negative in every condition.
Figure 2. Partial correlations controlling for dataset identity, across six prompting conditions. DSF (blue) is positive everywhere; three text-memorization metrics (orange/pink) are not. Bars show 95% CIs.
Result In zero-shot evaluation, consensus DSF is positively associated with accuracy (partial r = +0.41, 95% CI [+0.15, +0.62], p = 0.003, N = 54 model–dataset cells), and it stays positive across all six prompting conditions, peaking under few-shot (r = +0.62). ROUGE-L, BERTScore, and embedding similarity show no positive signal once dataset is controlled for (r = −0.19, −0.15, −0.16). The effect replicates on non-toxicity tasks: irony and subjectivity detection show partial r = +0.34 (N = 18).

Models whose elicited concept matches a dataset's definition tend to perform better on that task: text memorization gives no comparable signal. DSF needs only a handful of prompts and no labeled data, making it a practical model-selection diagnostic before committing to a full annotation run.

RQ2

Prompting rarely rescues a confident zero-shot error

To what extent can definitions, examples, or automated optimization correct zero-shot errors, and how does confidence relate to correctability?

We define the Rescue Rate as P(correct | prompted, zero-shot wrong): the probability that a better prompt fixes a mistake the model already made. Across 9 models, 5 datasets, and every non-trivial prompting condition (aligned definitions, few-shot examples, both combined, and two automated DSPy optimizers), the overall rescue rate is 34.8%. Almost two-thirds of zero-shot errors survive every intervention we tried, including a three-turn iterative reconsideration prompt (rescue rises from 7.5% at turn 1 to just 18.7% at turn 3).

Chart of rescue probability versus zero-shot confidence, showing an inverted-U shape. Rescue probability rises from moderate levels at very low confidence, peaks around 52% at 0.6-0.7 confidence, then falls sharply to about 21% for confidence above 0.9.
Figure 3. Rescue probability vs. zero-shot confidence, for zero-shot errors. Rescue peaks at 51.8% around 0.6–0.7 confidence, then drops to 20.8% above 0.9. The far-left drop is a distinct failure mode: short, context-deficient fragments (median 6 words) rather than confident mistakes.
Sticky Decision stickiness: each standard-deviation increase in zero-shot confidence cuts the odds of rescue by 16% (OR = 0.84, 95% CI [0.82, 0.87]). A correct zero-shot answer, conversely, predicts 6.4× higher odds of staying correct under prompting (OR = 6.43). Prompting mostly consolidates what the model already believes.
Rescue rate by model, aggregated across all prompted conditions
ModelRescue rate95% CIN (ZS errors)
Mistral-Small-24B44.2%[43.1, 45.2]8,536
DeepSeek-V338.1%[36.9, 39.2]7,264
Llama-3.1-8B37.5%[36.5, 38.5]9,040
GPT-4o-mini35.7%[34.6, 36.8]7,152
Mixtral-8x7B34.6%[33.5, 35.7]7,416
Mistral-7B31.6%[30.5, 32.6]7,568
Llama-3.1-70B31.3%[30.3, 32.4]7,848
Llama-3.3-70B30.4%[29.4, 31.4]7,592
Qwen-2.5-72B27.8%[26.8, 28.9]6,504

Steerability isn't a function of scale: Mistral-Small-24B out-rescues Llama-3.1-70B despite being a third of the size. “Bigger model” is not a substitute for “better matched model.”

RQ3

Wrong definitions shift verdicts, not confidence

How do LLMs behave under misaligned task definitions, and can confidence scores detect the mismatch?

We swap in definitions from a different dataset, e.g. asking a general-toxicity model to instead apply a hate-speech definition that requires identity-based targeting. Six such swaps span narrow (identity-based only) to broad (any disruptive content) scope. Models are clearly responsive to scope: narrow definitions shift predictions toward under-calling (bias −7 to −12 points), broad definitions toward over-calling (bias +9 to +13 points). They are not ignoring the instruction: they're following it, even when it's the wrong one.

Misalignment analysis, ordered narrow → broad definition scope. Confidence stays 85–91% in every row.
ConditionAcc.ChangeBiasConf.
Aligned definition82.011.3ref88.1
Fox News hate speech narrow82.617.5−7.290.6
Twitter hate speech narrow78.523.1−12.188.1
GameTox toxicity medium79.615.4+0.586.5
OLID offensive medium81.910.9+5.189.1
General toxicity broad81.19.2+8.885.0
Gaming toxicity broad76.413.3+13.188.5
Calibration curve chart plotting reported confidence against actual accuracy for zero-shot, aligned, and misaligned conditions. All three lines fall below the diagonal, indicating overconfidence, and are not clearly separated from each other.
Figure 4. Calibration curves: reported confidence vs. actual accuracy. Every condition sits below the diagonal (overconfident), and aligned vs. misaligned conditions are not meaningfully separated.
Miscalibrated Critical calibration failure: mean confidence barely moves: 87.0% zero-shot, 85.0–90.6% across aligned and misaligned definitions. The narrowest, most restrictive definition (Fox News hate speech) produces the highest confidence, not the lowest. Confidence reflects certainty given the instructions, not whether the instructions are the right ones.

The worst misaligned condition (Gaming toxicity, 76.4%) sits 5.6 points below aligned; the best (Fox News hate speech, 82.6%) actually outperforms the aligned definition. Across all models and datasets, definition wording alone produces up to 17% accuracy variation, more than double the ~5% variation from switching models. And steerability cuts both ways: Mistral-Small-24B has the highest misaligned rescue rate (73.7%) and the highest corruption rate (21.0%), the same responsiveness that lets a good definition help also lets a bad one hurt.

Definition design is a first-order experimental factor, not a fixed background assumption. And confidence thresholds cannot substitute for validating that the definition given to the model is the one you actually meant.

Try it: same message, seven definitions

Real numbers from Table 5. Click a definition, or let it auto-cycle: watch the verdict flip while confidence barely moves.

“LOL no you are just a jerk”

Definition in prompt

No definition provided.

Verbalized confidence

Confidence stays 85–91% in every condition, including the wrong ones.

Three lightweight safeguards for annotation pipelines

01

Measure definition alignment first

Run a DSF-style check before committing to a large-scale annotation job. It needs only a handful of prompts and no labeled data, and it flags which model–definition pairs are unlikely to work well.

02

Stress-test the definition, not just the model

Evaluate multiple plausible phrasings of the same policy and report sensitivity, meaning changes in bias, rescue, and corruption, instead of assuming one wording is robust.

03

Don't trust confidence as a policy check

A high confidence score means the model is sure given the instructions, not that the instructions match your intended policy. Use it to gauge certainty, never definition correctness.

Models & datasets

Nine instruction-tuned models (six core, three extended) evaluated on five primary toxicity datasets spanning social media, gaming, news, and forums, plus three supplementary datasets for robustness and generalization checks. 1,000 sampled instances per model–dataset–condition pair, temperature 0, across 10 manual prompting conditions plus 2 automated DSPy conditions.

Datasets
DatasetDomainSize% Positive
Twitter HateSocial media24,7835.8
OLIDSocial media14,10033.2
GameToxGaming53,00042.5
Fox NewsNews1,52828.5
Jigsaw Toxic CommentsForum159,5719.6
Jigsaw Unintended Bias †Forum97,3208.0
SemEval-2018 Irony †Social media4,61848.1
Subjectivity †Reviews10,00050.0

† Supplementary: robustness & generalization checks.

Models
ModelArch.Params
Llama-3.1-8BDense8B
Llama-3.1-70BDense70B
Mistral-7BDense7B
Mistral-Small-24BDense24B
Mixtral-8x7BMoE47B (13B act.)
DeepSeek-V3MoE671B (37B act.)
Llama-3.3-70BDense70B
GPT-4o-miniProp.~8B
Qwen-2.5-72BDense72B

BibTeX

@inproceedings{casanova2026limits,
  title     = {On the Limits of {LLM} Adaptability: Impact of
               Model-Internalized Priors on Annotation Task Performance},
  author    = {Casanova, Etienne and Kocielnik, Rafal and Alvarez, R. Michael},
  booktitle = {Proceedings of the 43rd International Conference on
               Machine Learning (ICML)},
  year      = {2026},
  note      = {Spotlight, Oral Presentation},
  url       = {https://openreview.net/forum?id=oTv2bKG5Qg}
}