Three node types
Document, Topic, and Entity nodes, joined by document–topic, document–entity, and entity–entity edges. The topic and entity nodes add the context that short headlines lack.
Short news headlines are hard to classify: they carry little context and labels are scarce. This work represents a news corpus as a graph of documents, topics, and entities, and classifies it semi-supervised with HGAT — enriching the graph's node features with pretrained word2vec embeddings to lift accuracy under few labels.
Text classification needs labeled data, which is expensive; short news headlines make it harder still, since a few words give little to learn from. Representing the corpus as a rich graph lets a small number of labels propagate to many unlabeled documents.
Each news document is connected to the topics it covers (from LDA) and the entities it mentions (from Wikipedia entity linking), forming a heterogeneous graph of three node types. HGAT embeds this graph with node- and type-level attention and classifies documents transductively. The contribution is at the node-feature stage: replacing sparse TF-IDF features with dense pretrained word2vec embeddings for document and entity nodes. On a 6,000-document AG News subset and a HuffPost News Category subset, this lifts accuracy over the TF-IDF-initialized baseline by 4.2 and 2.8 points. A released re-implementation re-runs the AG News comparison controlled — identical graph, splits, and seeds, averaged over five seeds — and confirms the effect: +4.9 points (paired p = 0.0011).
Raw news text is turned into a heterogeneous graph, each node type is given features, and HGAT embeds and classifies it. The pieces:
Document, Topic, and Entity nodes, joined by document–topic, document–entity, and entity–entity edges. The topic and entity nodes add the context that short headlines lack.
Entity nodes use Google-News word2vec vectors and document nodes a centroid of their word vectors — dense, low-dimensional features in place of sparse TF-IDF, keeping the rest of HGAT unchanged.
HGAT aggregates neighbors with both node-level attention (which neighbor matters) and type-level attention (which node type matters), so document, topic, and entity signals are weighted appropriately.
Entities come from TAGME (Wikipedia entity linking); topics come from LDA. Two documents are related when they share entities or topics, forming the meta-paths the model reasons over.
The whole corpus graph is available during training; only a few labels per class are used, and test documents provide unlabeled structure. Two datasets: an AG News subset and a HuffPost News Category subset.
Results come from one seed and split. Transductive GNNs at small label budgets vary a few points across seeds, so repeating over multiple seeds is the natural way to firm up the numbers.
The animation follows a single document through the model: its heterogeneous neighbourhood forms, node-level attention weighs each neighbour, type-level attention weighs each node type, and the aggregated representation is classified.
Holding the HGAT graph and architecture fixed, pretrained-embedding node features improve accuracy over the TF-IDF-initialized baseline on both datasets — +4.2 points on AG News and +2.8 on HuffPost, under a few labels per class. The released implementation re-runs the AG News comparison under a controlled multi-seed protocol and confirms it.
| Dataset | Node features | Accuracy (%) |
|---|---|---|
| AG News | TF-IDF (HGAT baseline) | 72.10 |
| AG News | word2vec (this work) | 76.30 |
| HuffPost | TF-IDF (baseline) | 56.75 |
| HuffPost | word2vec (this work) | 59.57 |
| Dataset | Node features | Accuracy (%) |
|---|---|---|
| AG News | TF-IDF (baseline) | 79.48 ± 1.19 |
| AG News | embeddings (this work — verified) | 84.38 ± 0.47 |
Gain: +4.90 points (paired t = 8.33, p = 0.0011) — the same finding as the 2021 single-run experiment, now established under a controlled protocol. Reproduce it with python run.py and python run.py --features tfidf from the repository's code/ folder.
A focused study: enrich the node features of a proven short-text graph model and measure the effect under label scarcity. What it sets out to do, what it does not, and where it can go next.
Short-text news classification is difficult because headlines are semantically sparse and labeled data are scarce. This paper does not propose a new architecture; it studies a single isolable design choice in HGAT: how the graph’s nodes are initialized. We keep the graph schema and dual-level attention fixed and replace high-dimensional TF-IDF node features with low-dimensional pretrained embeddings for entity and document nodes.
On a 6,000-document AG News subset and a HuffPost News Category subset, the embedding-initialized model reaches 76.30% and 59.57% accuracy in a transductive semi-supervised setting, improving over the TF-IDF-initialized HGAT by 4.20 and 2.82 points. Results are reported as single-run measurements, with a dedicated limitations section, and are confirmed by a controlled multi-seed re-run in the released implementation. The paper refines the author’s earlier submitted thesis document — corrected method description, accurate dataset naming, and verified results.
The paper opens in the browser; the code and dataset open on GitHub.