Caption Distillation with Uncertainty Modeling for Local Visual Memory and Semantic Retention in Private Image Archives

Main Article Content

Ahmed Salem
Yahya Mohamed
Ely Cheikh

Abstract

Personal image archives now contain thousands of photographs whose utility depends less on global object recognition than on remembering small visual distinctions across time. A private visual memory system must index these archives without sending images to a remote service, while still producing captions that are specific enough for later search, clustering, and reminder queries. This paper studies uncertainty aware caption distillation for on device visual memory. The proposed method compresses several candidate captions for an image into a short memory record that preserves stable entities, attributes, locations, and relations while suppressing hallucinated or low confidence details. Candidate captions are produced by a compact vision language model, scored by calibrated token level uncertainty, and distilled into a structured sentence using a local language decoder constrained by visual evidence. We construct three evaluation sets: HomeArchive-18K, EventRecall-9K, and ObjectTrace-6K. The method reduces memory record length by 43.2% relative to dense captions while improving later retrieval mean reciprocal rank by 7.9% over a compact caption baseline. On a simulated private assistant task, it lowers false visual reminders from 14.8\% to 8.6\% and maintains 91.3% of attribute recall under an 80 byte average caption budget. Ablations suggest that entropy based token rejection and cross caption agreement are both necessary when models are quantized for mobile inference. The findings indicate that small on device models can support useful visual memory when caption generation is treated as selective evidence retention rather than fluent description alone.

Article Details

Section

Articles

References