Conceptual

Limits of LLM Self-Refinement for Product Attribute Value Extraction

An empirical finding that two automated self-refinement techniques for LLMs — error-based prompt rewriting and post-hoc self-correction — fail to significantly improve product attribute value extraction across zero-shot, few-shot, and fine-tuning settings with GPT-4o, while substantially increasing token and processing cost. When development data is available, fine-tuning yields the best accuracy and amortizes its ramp-up cost at scale.