HGAT · Short-Text News
Research paper  ·  Sujil Devkota  ·  2021

A heterogeneous graph attention network for short-text news classification.

Short news headlines are hard to classify: they carry little context and labels are scarce. This work represents a news corpus as a graph of documents, topics, and entities, and classifies it semi-supervised with HGAT — enriching the graph's node features with pretrained word2vec embeddings to lift accuracy under few labels.

Overview

The approach, in short.

Text classification needs labeled data, which is expensive; short news headlines make it harder still, since a few words give little to learn from. Representing the corpus as a rich graph lets a small number of labels propagate to many unlabeled documents.

■ In one paragraph

Each news document is connected to the topics it covers (from LDA) and the entities it mentions (from Wikipedia entity linking), forming a heterogeneous graph of three node types. HGAT embeds this graph with node- and type-level attention and classifies documents transductively. The contribution is at the node-feature stage: replacing sparse TF-IDF features with dense pretrained word2vec embeddings for document and entity nodes. On a 6,000-document AG News subset and a HuffPost News Category subset, this lifts accuracy over the TF-IDF-initialized baseline by 4.2 and 2.8 points. A released re-implementation re-runs the AG News comparison controlled — identical graph, splits, and seeds, averaged over five seeds — and confirms the effect: +4.9 points (paired p = 0.0011).

The idea
Richer node features
Pretrained embeddings on a fixed HGAT graph.
The setting
Semi-supervised
Transductive, few labels per class.
The result
+4.9 points, verified
Controlled 5-seed re-run (2021 single-run: +4.2 / +2.8).
84.4%
AG News accuracy with embedding features — verified re-run, 5 seeds
+4.9
Points over the TF-IDF baseline, controlled (p = 0.0011)
3
Node types — Document · Topic · Entity
6,000
Documents per corpus, few labels per class
Method at a glance

How the pipeline is put together.

Raw news text is turned into a heterogeneous graph, each node type is given features, and HGAT embeds and classifies it. The pieces:

GraphD · T · E

Three node types

Document, Topic, and Entity nodes, joined by document–topic, document–entity, and entity–entity edges. The topic and entity nodes add the context that short headlines lack.

Contributionnode features

Pretrained word2vec features

Entity nodes use Google-News word2vec vectors and document nodes a centroid of their word vectors — dense, low-dimensional features in place of sparse TF-IDF, keeping the rest of HGAT unchanged.

Modelattention

Dual-level attention

HGAT aggregates neighbors with both node-level attention (which neighbor matters) and type-level attention (which node type matters), so document, topic, and entity signals are weighted appropriately.

Signalsentities · topics

Entities and topics

Entities come from TAGME (Wikipedia entity linking); topics come from LDA. Two documents are related when they share entities or topics, forming the meta-paths the model reasons over.

Setuptransductive

Semi-supervised training

The whole corpus graph is available during training; only a few labels per class are used, and test documents provide unlabeled structure. Two datasets: an AG News subset and a HuffPost News Category subset.

Limitationvariance

Single-run results

Results come from one seed and split. Transductive GNNs at small label budgets vary a few points across seeds, so repeating over multiple seeds is the natural way to firm up the numbers.

Architecture

How one document gets classified.

The animation follows a single document through the model: its heterogeneous neighbourhood forms, node-level attention weighs each neighbour, type-level attention weighs each node type, and the aggregated representation is classified.

Results

The gain from richer node features.

Holding the HGAT graph and architecture fixed, pretrained-embedding node features improve accuracy over the TF-IDF-initialized baseline on both datasets — +4.2 points on AG News and +2.8 on HuffPost, under a few labels per class. The released implementation re-runs the AG News comparison under a controlled multi-seed protocol and confirms it.

Test accuracy (%) with pretrained-embedding vs. TF-IDF node features on the same graph. Highlighted rows are this work; single run.
DatasetNode featuresAccuracy (%)
AG NewsTF-IDF (HGAT baseline)72.10
AG Newsword2vec (this work)76.30
HuffPostTF-IDF (baseline)56.75
HuffPostword2vec (this work)59.57
Verified re-run in the released implementation: identical graph, splits, and seeds — only the node features change. Averaged over 5 seeds (± unbiased std). Entity linking and word vectors use offline substitutes (spaCy NER + GloVe), disclosed in the repository.
DatasetNode featuresAccuracy (%)
AG NewsTF-IDF (baseline)79.48 ± 1.19
AG Newsembeddings (this work — verified)84.38 ± 0.47

Gain: +4.90 points (paired t = 8.33, p = 0.0011) — the same finding as the 2021 single-run experiment, now established under a controlled protocol. Reproduce it with python run.py and python run.py --features tfidf from the repository's code/ folder.

Contributions & scope

What the work delivers, and its bounds.

A focused study: enrich the node features of a proven short-text graph model and measure the effect under label scarcity. What it sets out to do, what it does not, and where it can go next.

Contributions

What the work delivers

  • Pretrained-embedding node features for HGAT
  • A document/topic/entity graph pipeline for news
  • A gain over the TF-IDF baseline on two datasets
  • Results reported at few labels per class
Scope

What it does not claim

  • No change to the HGAT architecture itself
  • Single-run results, not multi-seed averages
  • Two news corpora, not a full benchmark suite
  • A transductive, not inductive, setting
Future work

Natural next steps

  • Averaging over seeds and more label budgets
  • Contextual embeddings for node features
  • A neural topic model in place of LDA
  • More short-text datasets and an inductive variant
The paper

Scoped to exactly what the evidence supports.

Abstract

Short-text news classification is difficult because headlines are semantically sparse and labeled data are scarce. This paper does not propose a new architecture; it studies a single isolable design choice in HGAT: how the graph’s nodes are initialized. We keep the graph schema and dual-level attention fixed and replace high-dimensional TF-IDF node features with low-dimensional pretrained embeddings for entity and document nodes.

On a 6,000-document AG News subset and a HuffPost News Category subset, the embedding-initialized model reaches 76.30% and 59.57% accuracy in a transductive semi-supervised setting, improving over the TF-IDF-initialized HGAT by 4.20 and 2.82 points. Results are reported as single-run measurements, with a dedicated limitations section, and are confirmed by a controlled multi-seed re-run in the released implementation. The paper refines the author’s earlier submitted thesis document — corrected method description, accurate dataset naming, and verified results.

  • 01Focused contribution. Node-feature initialization; the HGAT architecture is credited to Linmei et al.
  • 02Complete method. The graph construction and dual-level attention are specified in full.
  • 03Clear reporting. Single-run results, with each row’s setup stated in the table.
  • 04Stated limits. A dedicated section covers what the setting does and does not establish.