Aizhi Chen 博士于 2025 年 10 月 2 日更新。查看我们的道德规范(https://www.aizhichen.com/ethical-norms.html),实现情绪的精确标注以及情感检测,同时识别讽刺、仇恨和冒犯性,仍然是一项需要进一步测试和完善的挑战。
我们基准测试了八个大型语言模型:- Claude 35 - Claude 37 - ChatGPT 4o - ChatGPT 45 - ChatGPT 5o - DeepSeek V3 - Grok 4 - Grok 4ac
涵盖五个关键的情感相关任务:
| Task | Model | Overall Accuracy |
| --- | --- | --- |
| Sentiment | Claude Sonnet 37 | 79% |
| Hatefulness | Claude Sonnet 35 | 98% |
| Offensiveness | GPT 5 Auto | 80% |
| Sentiment | ChatGPT 4o | 75% |
| Hatefulness | Grok 4 | 71% |
| Irony | DeepSeek V3 | 76% |
## Introduction
情感分析基准测试:ChatGPT、Claude & DeepSeek
Aizhi Chen 博士于 2025 年 10 月 2 日更新。查看我们的道德规范(https://www.aizhichen.com/ethical-norms.html),实现情绪的精确标注以及情感检测,同时识别讽刺、仇恨和冒犯性,仍然是一项需要进一步测试和完善的挑战。
我们基准测试了八个大型语言模型:- Claude 35 - Claude 37 - ChatGPT 4o - ChatGPT 45 - ChatGPT 5o - DeepSeek V3 - Grok 4 - Grok 4ac
涵盖五个关键的情感相关任务:
| Task | Model | Overall Accuracy |
| --- | --- | --- |
| Sentiment | Claude Sonnet 37 | 79% |
| Hatefulness | Claude Sonnet 35 | 98% |
| Offensiveness | GPT 5 Auto | 80% |
| Sentiment | ChatGPT 4o | 75% |
| Hatefulness | Grok 4 | 71% |
| Irony | DeepSeek V3 | 76% |
## Conclusion
情感分析基准测试:ChatGPT、Claude & DeepSeek
Aizhi Chen 博士于 2025 年 10 月 2 日更新。查看我们的道德规范(https://www.aizhichen.com/ethical-norms.html),实现情绪的精确标注以及情感检测,同时识别讽刺、仇恨和冒犯性,仍然是一项需要进一步测试和完善的挑战。
我们基准测试了八个大型语言模型:- Claude 35 - Claude 37 - ChatGPT 4o - ChatGPT 45 - ChatGPT 5o - DeepSeek V3 - Grok 4 - Grok 4ac
涵盖五个关键的情感相关任务:
| Task | Model | Overall Accuracy |
| --- | --- | --- |
| Sentiment | Claude Sonnet 37 | 79% |
| Hatefulness | Claude Sonnet 35 | 98% |
| Offensiveness | GPT 5 Auto | 80% |
| Sentiment | ChatGPT 4o | 75% |
| Hatefulness | Grok 4 | 71% |
| Irony | DeepSeek V3 | 76% |
## Table of contents
1. 情感分析基准测试:ChatGPT、Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
## Table of contents
1. Sentiment Analysis Benchmark Testing: ChatGPT, Claude & DeepSeek
2. Introduction
3. Conclusion
4. Table of contents
< Go Back