Introduction

This report evaluates the performance of two annotation strategies — human-generated and LLM-generated text annotations — for classifying news images into five predefined categories: Arts, Business, Technology, U.S., and World. The goal is to assess the effectiveness of each approach in supporting automated news section classification.

Methodology

Two DistilBERT-based multi-class classification models were fine-tuned:

  • LLM-Annotated Model: Trained on captions generated by a large language model based on 250 new images.
  • Human-Annotated Model: Trained on captions written by human annotators based on the same images.

Each model was evaluated using standard classification metrics: accuracy, macro-averaged F1-score, and per-class F1-scores. Confusion matrices were also analyzed to assess misclassification patterns.

Findings

As for the overall performance, the human-annotated model is more than twice as accurate and balanced in its predictions. Accuracy is the percentage of correct predictions, while Macro F1 averages F1-scores across all classes equally, showing how well the model performs across the board. By the number, the human annotated model outperforms.

Figure 1: Overall Accuracy and Macro F1-score

As per-class in the F1-Scores, the human model maintains strong performance across all categories (eg. US, World) ranging from 0.6 to 0.9. The lowest score is on U.S. (F1: 0.60), which is still better than the LLM’s best. The LLM model performed best on Technology (F1: 0.43) but poorly on Arts (F1: 0.17) and Business (F1: 0.20). It suggests that the LLM model is inconsistent and underperforms across all categories.

Figure 2: F1-score per Category

The pattern of mis-prediction is assessed as well: In general, the human model’s confusion matrix is mostly diagonal, meaning it correctly predicts the actual category most of the time. However, the LLM model’s confusion matrix shows many misclassifications, especially between Technology, U.S., and World. This may suggest that the model often predicts the wrong category.

Figure 3: Confusion Matrix Heatmaps

One of the possible reasons that the human annotation outperforms LLM’s one is that human annotations are typically more context-aware, nuanced, and tailored to the visual and journalistic content of the image. This may explain why it gets the highest f1-score on Arts in figure 2. By contrast, LLM-generated annotations rely heavily on the prompt, which may be too generic or lack the specificity to distinguish between similar news categories (e.g., U.S. vs. World) in figure 3.

Conclusion

The result of the two models suggests that human annotations lead to significantly better model performance in news image classification tasks, particularly on context-specific issues and on distinguishing between similar topics. However, the result of llm-model is heavily based on the prompt generating annotation from the images. It is hard to conclude that LLM’s model can by no means outperform human annotation.

Collab link: https://colab.research.google.com/drive/1jmBEruHApg1FJQcLS56F-EPRq62tV1VU?usp=sharing

Skye Tang Avatar

Published by

Categories:

Leave a Reply

Discover more from Skye Tang's Portfolio

Subscribe now to keep reading and get access to the full archive.

Continue reading