Lab Work 4: Text Classification Using Transfer Learning

Soal / Brief

Write a text classification program with the following specifications:

  1. Use the dataset: https://bit.ly/LabWork3KitIF5153
  2. Compare and analyze the tokenization results from several tokenizers, relating them to the type of tokenizer (min: WordPiece, BPE, SentencePiece)
  3. Perform fine-tuning on at least three Encoder models (min: BERT, RoBERTa, XLM-R), analyze and compare the results of all three, as well as with the results from previous Shallow Learning and Deep Learning approaches!

Submit to Kaggle: https://www.kaggle.com/competitions/if-5153-advanced-nlp-lab-work-3

  • You must have at least one submission before the class ends (11 AM that day).
  • Change the Team Name on the Kaggle competition to your Student ID (NIM).

Create an analysis report containing:

  • An explanation of the reasoning behind the chosen preprocessing techniques.
  • A comparison of the tokenization results across the different tokenizers used (WordPiece, BPE, SentencePiece), and how each tokenizer’s characteristics relate to the fine-tuning performance of its corresponding Encoder model.
  • A comparison of the results from the three (or more) fine-tuned Encoder models, and their impact on model performance relative to each other and to the previous Shallow Learning (Lab Work 1) and Deep Learning (Lab Work 2) approaches.
  • Insights gained.

Ketentuan / Yang Dikumpulkan

  • Notebook mengikuti struktur template resmi (Lab Work 4 Template IF5153 NLP.ipynb), semua TODO diisi
  • Minimal 1 submission Kaggle sebelum kelas berakhir (sudah, dari versi bert-only)
  • Rename Team Name Kaggle jadi NIM (13523057)
  • Fine-tuning 3 encoder model: BERT, RoBERTa, XLM-R
  • Perbandingan 3 tokenizer (WordPiece/BPE/SentencePiece)
  • Laporan 04-13523057-report.pdf (belum ditulis — 4 bagian wajib: preprocessing reasoning, tokenizer comparison, model comparison vs Lab 1/2, insights)
  • Finalisasi 04-13523057-notebook.ipynb (pilih 1 dari beberapa varian yang sudah dicoba)
  • Zip jadi 04-13523057.zip
  • Deadline: Minggu, 27 September 2026, 23:59

Progress & Catatan Pengerjaan

Status per 2026-09-26: submission terbaik di Kaggle sejauh ini public score 0.88160 (ensemble “stacked” dari BERT+RoBERTa+XLM-R, full data, 3 epoch, class-weighted loss), dari notebook 04-13523057-notebook-maksimal.ipynb / hasil run notebookaf3ce78c84.ipynb.

Eksperimen yang sudah dicoba (real run di Kaggle, bukan cuma smoke test):

  • BERT-only: 0.87557
  • Ensemble 3-strategi (simple/weighted/stacked): 0.88160 (stacked menang) — skor terbaik
  • Ensemble + “optimized weights” versi buggy (overfit ke validation set): 0.88019 (lebih rendah — ketauan dari selisih val-score vs public-score, sudah diperbaiki jadi evaluasi fair pakai holdout 30%)
  • Diverse training subsample per model (bagging-style): val macro F1 turun konsisten ~0.004-0.005 di semua model & strategi ensemble dibanding versi full-data — tidak disubmit ke Kaggle karena sudah jelas lebih rendah di validation. Insight: diversitas arsitektur (BERT/RoBERTa/XLM-R beda tokenizer & pretraining corpus) sudah cukup untuk ensembling; mengorbankan 10% data training demi variasi tambahan tidak worth it untuk model pretrained besar (beda dengan bagging pada model high-variance seperti decision tree).
  • Focal loss (gamma=2) — sedang dijalankan user di Kaggle, belum ada hasil.

Next steps:

  1. Tunggu hasil focal loss run.
  2. Mulai tulis 04-13523057-report.pdf, pakai konten dari notebookaf3ce78c84.ipynb (hasil submission terbaik) sebagai sumber utama data/angka:
    • Preprocessing reasoning: cell 15 (mojibake fix + casing per model, alasan skip stemming/stopword removal karena subword tokenizer sudah handle itu sendiri).
    • Tokenizer comparison: cell 18-22 (vocab size, avg token/row: BERT 17.6, RoBERTa/XLM-R ~19.2; OOV handling table; sample tokenization side-by-side).
    • Model comparison: cell 35 (tabel F1 BERT 0.8742 > RoBERTa 0.8631 > XLM-R 0.8620) + confusion matrix cell 36 + dibandingkan ke Lab Work 2 (LSTM+GloVe 0.8083, ensemble Lab 2 0.8404 — semua kalah dari BERT sendirian). Nomor Lab Work 1 (shallow ML) masih perlu dicari.
    • Insights: kenapa BERT menang meski RoBERTa “lebih kuat” secara umum (kemungkinan: teks pendek/informal + WordPiece cocok untuk domain slang campuran ini), kenapa ensembling (stacked) masih menang tipis dari BERT sendirian, dan insight dari eksperimen diversesplit (di atas).
  3. Setelah dapat hasil focal loss, putuskan submission final mana yang dipakai untuk notebook resmi.
  4. Finalisasi & zip deliverables.

Files Terkait

  • Kerja utama ada di luar vault: D:\Kuliah\Obsidian-Brainverse\_Materi\NLP\tugas-lab\Labwork4\ (dataset, build scripts .ps1, semua varian notebook, hasil run Kaggle yang didownload)
  • Brief asli: _Materi/NLP/tugas-lab/Labwork4/Lab Work #4 Transfer Learning-Text Classification.pptx.pdf
  • Template resmi: _Materi/NLP/tugas-lab/Labwork4/Lab Work 4 Template IF5153 NLP.ipynb
  • Notebook sumber submission terbaik (0.88160): _Materi/NLP/tugas-lab/Labwork4/notebookaf3ce78c84.ipynb
  • Report Lab Work 2 (untuk angka benchmark): _Materi/NLP/tugas-lab/02-13523057-report.pdf