Взаємні самооцінки загальних характеристик деяких популярних безоплатних інструментів ШІ
Loading...
Date
item.page.thesis.degree.name
item.page.thesis.degree.level
item.page.thesis.degree.discipline
item.page.thesis.degree.department
item.page.thesis.degree.grantor
item.page.thesis.degree.advisor
item.page.thesis.degree.committeeMember
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
На теперішній час з’являється велика кількість публікацій, присвячених аналізу, оцінці та оптимальному вибору моделей ШІ для конкретних застосувань, і кількість ринкових пропози цій щодо надання послуг по порівнянню і вибору інструментів ШІ для конкретних задач. Поряд із цим інформаційним мейнстрімом заслуговує на інтерес огляд можливостей і характеристик ін струментів на основі «власних думок» самих моделей ШІ відносно себе і своїх «колег». Мета статті полягає в експериментальному порівняльному огляді оцінок характеристик моделей і ви явленні можливих тенденцій у цих оцінках. Для огляду обрані поточні покоління популярних досту пних інтелектуальних асистентів, орієнтованих на широке коло користувачів і їх задачі. Зокрема, одержуються й обговорюються взаємооцінки характеристик моделей ChatGPT версій 3.5+DALL E,4o і 5, DeepSeek V3, Gemini 1.5 і 2.5, Claude Sonnet 3 і 4). Наводяться й аналізуються перехресні 5-бальні самооцінки поточних поколінь моделей за заданими критеріями та їх оцінки очікуваних можливостей ChatGPT-5. Обговорюються переваги моделей у контексті збалансованих оцінок та рекомендацій щодо їх застосування. Виявлено неоднозначність, суперечливість і можливі відхи лення від об’єктивності. Моделі часто демонструють упередженість, завищуючи власні можли вості. Серед можливих причин цього (застарілість деяких баз знань, своєрідна галюціногенність тощо) найбільш імовірною вважається конкуренто-маркетингова спрямованість моделей різних виробників. Тому при використанні самих моделей для відповідних оцінок і вибору суттєвою має бути саме зважена колективність їх думок і виключення самооцінок. Відзначається значення поя ви GPT-5 у змінах «конкурентного ландшафту» і підходах до критеріїв оцінок моделей.
Currently, there is a growing number of publications devoted to the analysis, evaluation, and optimal selection of AI models for specific applications, as well as market offers for services comparing and selecting AI tools for specific tasks. Along with this mainstream information, it is interesting to review the capabilities and characteristics of tools based on AI models’ «opinions» about themselves and their «colleagues». The purpose of this article is to conduct an experimental comparative review of model characteristic assessments and identify possible trends in them. The current generation of popular, acces sible intelligent assistants, aimed at a wide range of users and their tasks, have been chosen for review. In particular, mutual assessments of the characteristics of ChatGPT versions 3.5+DALL-E, 4o, and 5, DeepSeek V3, Gemini 1.5 and 2.5, Claude Sonnet 3 and 4 are obtained and discussed. Cross-referenced 5-point self-assessments of current model generations according to specified criteria and their assess ments of the expected capabilities of ChatGPT-5 are provided and analyzed. The advantages of models are discussed in the context of balanced assessments and recommendations for their application. Ambigui ty, inconsistency, and possible deviations from objectivity have been identified. Models often demonstrate bias by overestimating their own capabilities. Among the possible reasons for this (obsolescence of some knowledge bases, some kind of hallucinogenicity, etc.), the most likely is considered to be the competitive and marketing orientation of models from different manufacturers. Therefore, when using the models themselves for relevant assessments and selection, it is essential to weigh their opinions collectively and exclude self-assessments. The emergence of GPT-5 is noted for its significance in changes to the «com petitive landscape» and approaches to model evaluation criteria.
Currently, there is a growing number of publications devoted to the analysis, evaluation, and optimal selection of AI models for specific applications, as well as market offers for services comparing and selecting AI tools for specific tasks. Along with this mainstream information, it is interesting to review the capabilities and characteristics of tools based on AI models’ «opinions» about themselves and their «colleagues». The purpose of this article is to conduct an experimental comparative review of model characteristic assessments and identify possible trends in them. The current generation of popular, acces sible intelligent assistants, aimed at a wide range of users and their tasks, have been chosen for review. In particular, mutual assessments of the characteristics of ChatGPT versions 3.5+DALL-E, 4o, and 5, DeepSeek V3, Gemini 1.5 and 2.5, Claude Sonnet 3 and 4 are obtained and discussed. Cross-referenced 5-point self-assessments of current model generations according to specified criteria and their assess ments of the expected capabilities of ChatGPT-5 are provided and analyzed. The advantages of models are discussed in the context of balanced assessments and recommendations for their application. Ambigui ty, inconsistency, and possible deviations from objectivity have been identified. Models often demonstrate bias by overestimating their own capabilities. Among the possible reasons for this (obsolescence of some knowledge bases, some kind of hallucinogenicity, etc.), the most likely is considered to be the competitive and marketing orientation of models from different manufacturers. Therefore, when using the models themselves for relevant assessments and selection, it is essential to weigh their opinions collectively and exclude self-assessments. The emergence of GPT-5 is noted for its significance in changes to the «com petitive landscape» and approaches to model evaluation criteria.
Description
Citation
Литвинов, В. А. Взаємні самооцінки загальних характеристик деяких популярних безоплатних інструментів ШІ / В. А. Литвинов, С. В. Грибков, І. М. Оксанич // Математичні машини і системи. – 2025. – № 3-4. – С. 3–12.
