Keywords
generative artificial intelligence, inferentiality, I2T generation, pragmatic analysis, referentiality
Abstract
The study aims to conduct a comparative analysis of ChatGPT, Grok, and Gemini to evaluate their capacities in image-to-text generation. The researcher utilized three AI-generated image prompts from Canva, which were processed through the mentioned generative AI (GenAI) models. With the aid of a standardized text prompt, each model was instructed to produce eighteen descriptive sentences corresponding to the prompts. The study examined the contextual alignment, pragmatic appropriateness, as well hallucinations and pragmatic errors present in the generated sentences through pragmatic analysis. The analysis revealed the use of appropriate referents, modal markers, and descriptive details that contributed to contextually and pragmatically accurate descriptions of the images. Despite this, several structural inconsistencies and lexical ambiguities surfaced. While the outputs demonstrated logical coherence through the recognition of lexical patterns, the findings indicate that the models heavily rely on probabilistic associations rather than genuine semantic understanding. The limitations of GenAI became evident in descriptions that contained hallucinations, excessive inference, and syntactic errors, particularly in cases where visual evidence was absent or insufficient to support the inferred content. Overall, Gemini outperformed the other models in producing accurate and meaningful descriptions with high contextual and pragmatic precision, followed by ChatGPT, whose outputs showed some structural flaws, Grok produced the highest number of hallucinated outputs and pragmatic errors.
Recommended Citation
Bosque, Ariel U.
(2026)
"Pragmatic Analysis of AI Image-to-Text Generation: The Case of ChatGPT, Grok, and Gemini,"
Social Sciences and Development Review: Vol. 18:
No.
1, Article 10.
Available at:
https://scholar.pup.edu.ph/ssdr/vol18/iss1/10







