Lab Work 4: Text Classification Using Transfer Learning
Soal / Brief
Write a text classification program with the following specifications:
- Use the dataset: https://bit.ly/LabWork3KitIF5153
- Compare and analyze the tokenization results from several tokenizers, relating them to the type of tokenizer (min: WordPiece, BPE, SentencePiece)
- Perform fine-tuning on at least three Encoder models (min: BERT, RoBERTa, XLM-R), analyze and compare the results of all three, as well as with the results from previous Shallow Learning and Deep Learning approaches!
Submit to Kaggle: https://www.kaggle.com/competitions/if-5153-advanced-nlp-lab-work-3
- You must have at least one submission before the class ends (11 AM that day).
- Change the Team Name on the Kaggle competition to your Student ID (NIM).
Create an analysis report containing:
- An explanation of the reasoning behind the chosen preprocessing techniques.
- A comparison of the tokenization results across the different tokenizers used (WordPiece, BPE, SentencePiece), and how each tokenizer’s characteristics relate to the fine-tuning performance of its corresponding Encoder model.
- A comparison of the results from the three (or more) fine-tuned Encoder models, and their impact on model performance relative to each other and to the previous Shallow Learning (Lab Work 1) and Deep Learning (Lab Work 2) approaches.
- Insights gained.
Ketentuan / Yang Dikumpulkan
- Notebook mengikuti struktur template resmi (
Lab Work 4 Template IF5153 NLP.ipynb), semua TODO diisi - Minimal 1 submission Kaggle sebelum kelas berakhir (sudah, dari versi bert-only)
- Rename Team Name Kaggle jadi NIM (13523057)
- Fine-tuning 3 encoder model: BERT, RoBERTa, XLM-R
- Perbandingan 3 tokenizer (WordPiece/BPE/SentencePiece)
- Laporan
04-13523057-report.pdf(belum ditulis — 4 bagian wajib: preprocessing reasoning, tokenizer comparison, model comparison vs Lab 1/2, insights) - Finalisasi
04-13523057-notebook.ipynb(pilih 1 dari beberapa varian yang sudah dicoba) - Zip jadi
04-13523057.zip - Deadline: Minggu, 27 September 2026, 23:59
Progress & Catatan Pengerjaan
Status per 2026-09-26: submission terbaik di Kaggle sejauh ini public score 0.88160 (ensemble “stacked” dari BERT+RoBERTa+XLM-R, full data, 3 epoch, class-weighted loss), dari notebook 04-13523057-notebook-maksimal.ipynb / hasil run notebookaf3ce78c84.ipynb.
Eksperimen yang sudah dicoba (real run di Kaggle, bukan cuma smoke test):
- BERT-only: 0.87557
- Ensemble 3-strategi (simple/weighted/stacked): 0.88160 (stacked menang) — skor terbaik
- Ensemble + “optimized weights” versi buggy (overfit ke validation set): 0.88019 (lebih rendah — ketauan dari selisih val-score vs public-score, sudah diperbaiki jadi evaluasi fair pakai holdout 30%)
- Diverse training subsample per model (bagging-style): val macro F1 turun konsisten ~0.004-0.005 di semua model & strategi ensemble dibanding versi full-data — tidak disubmit ke Kaggle karena sudah jelas lebih rendah di validation. Insight: diversitas arsitektur (BERT/RoBERTa/XLM-R beda tokenizer & pretraining corpus) sudah cukup untuk ensembling; mengorbankan 10% data training demi variasi tambahan tidak worth it untuk model pretrained besar (beda dengan bagging pada model high-variance seperti decision tree).
- Focal loss (gamma=2) — sedang dijalankan user di Kaggle, belum ada hasil.
Next steps:
- Tunggu hasil focal loss run.
- Mulai tulis
04-13523057-report.pdf, pakai konten darinotebookaf3ce78c84.ipynb(hasil submission terbaik) sebagai sumber utama data/angka:- Preprocessing reasoning: cell 15 (mojibake fix + casing per model, alasan skip stemming/stopword removal karena subword tokenizer sudah handle itu sendiri).
- Tokenizer comparison: cell 18-22 (vocab size, avg token/row: BERT 17.6, RoBERTa/XLM-R ~19.2; OOV handling table; sample tokenization side-by-side).
- Model comparison: cell 35 (tabel F1 BERT 0.8742 > RoBERTa 0.8631 > XLM-R 0.8620) + confusion matrix cell 36 + dibandingkan ke Lab Work 2 (LSTM+GloVe 0.8083, ensemble Lab 2 0.8404 — semua kalah dari BERT sendirian). Nomor Lab Work 1 (shallow ML) masih perlu dicari.
- Insights: kenapa BERT menang meski RoBERTa “lebih kuat” secara umum (kemungkinan: teks pendek/informal + WordPiece cocok untuk domain slang campuran ini), kenapa ensembling (stacked) masih menang tipis dari BERT sendirian, dan insight dari eksperimen diversesplit (di atas).
- Setelah dapat hasil focal loss, putuskan submission final mana yang dipakai untuk notebook resmi.
- Finalisasi & zip deliverables.
Files Terkait
- Kerja utama ada di luar vault:
D:\Kuliah\Obsidian-Brainverse\_Materi\NLP\tugas-lab\Labwork4\(dataset, build scripts.ps1, semua varian notebook, hasil run Kaggle yang didownload) - Brief asli:
_Materi/NLP/tugas-lab/Labwork4/Lab Work #4 Transfer Learning-Text Classification.pptx.pdf - Template resmi:
_Materi/NLP/tugas-lab/Labwork4/Lab Work 4 Template IF5153 NLP.ipynb - Notebook sumber submission terbaik (0.88160):
_Materi/NLP/tugas-lab/Labwork4/notebookaf3ce78c84.ipynb - Report Lab Work 2 (untuk angka benchmark):
_Materi/NLP/tugas-lab/02-13523057-report.pdf