Exercise for IF5153 Advanced Natural Language Processing — Topic: Sentence Comparison
Soal / Brief
(5th October 2026)
1. BGE-M3 Hybrid Retrieval
Query: “laptop for programming” (3 tokens: laptop, for, programming) Document 1: “best notebook computer for software developers” Document 2: “laptop sleeve and carrying case for programming books”
1.1 Prediction
- Before doing any calculations: based on your intuition, which document is more relevant to the query?
- Of the three BGE-M3 retrieval methods (dense, sparse, and multi-vector), which one is most likely to select the wrong document?
- Briefly explain your reasoning, and compare your prediction with your calculation results in 1.6.
1.2 Dense Score — Given dense vectors (simplified to 3 dimensions): , , . Calculate for D1 and D2.
1.3 Sparse Score — Token weights (Linear layer output ), only tokens appearing in both query and document:
| Token | |||
|---|---|---|---|
| laptop | 0.50 | tidak muncul | 0.40 |
| for | 0.05 | 0.04 | 0.05 |
| programming | 0.45 | tidak muncul | 0.30 |
Calculate for D1 and D2. Note: “laptop” and “notebook” are different tokens.
1.4 Multi-vector Score — Cosine similarity between each query token (row) and each D1 token (column):
| Query \ D1 | best | notebook | computer | for | software | developers |
|---|---|---|---|---|---|---|
| laptop | 0.35 | 0.93 | 0.80 | 0.10 | 0.20 | 0.15 |
| for | 0.08 | 0.05 | 0.06 | 0.99 | 0.04 | 0.07 |
| programming | 0.12 | 0.18 | 0.40 | 0.06 | 0.90 | 0.85 |
Calculate . For D2, max similarity per query token already given: laptop=0.90, for=0.95, programming=0.85. Calculate .
1.5 Join — With (dense), (sparse), (multi-vec), calculate for D1 and D2. Which document ranks first?
1.6 Analysis
- If only one retrieval method were used (dense/sparse/multi-vector only), which document ranks first for each? Which method is “fooled” by lexical overlap?
- Keep . At what value of does D2 begin to outrank D1? (Hint: find such that .) What does this tell us about the importance of selecting appropriate weights for hybrid retrieval? Why does the BGE-M3 paper initially assign a relatively small weight to sparse retrieval?
- Give one other example of an English query and document where sparse retrieval would outperform dense retrieval (hint: product names, error codes, highly specific technical terms).
2. BGE-M3 Paper Exploration
Read the original paper at arXiv 2402.03216 (or its HTML version). For each answer, include the page number and/or section where you found the information.
a. Data Curation — Total text pairs used for unsupervised pre-training + number of languages covered? Name ≥3 sources of unsupervised data. What was the synthetic data generated using GPT-3.5 intended for, and why was available real-world data insufficient for this purpose?
b. Efficient Batching — What is length-grouped sampling, and what problem does it solve? What is split-batch with gradient checkpointing, and by what factor does the paper claim it can increase batch size for inputs of 8,192 tokens?
c. Self-Knowledge Distillation — Find the initial values of (dense) and (sparse) used at the beginning of the fine-tuning stage. Why are the weights not set equally from the beginning? Relate to the quality of the sparse retrieval component at the start of training.
d. Independent Exploration: Beyond BGE-M3 — Choose ONE other model described as a “Variation on SBERT” not covered in lecture slides (lecture covered SimCSE, E5, BGE-M3). Find the original paper and summarize in 3-4 sentences: (1) main contribution, (2) how it differs from basic SBERT, (3) link to the paper.
Ketentuan / Yang Dikumpulkan
- 1.1–1.6: prediksi + perhitungan dense/sparse/multi-vector/join score + analisis bobot
- 2a–2c: jawaban dari paper BGE-M3 asli (arXiv 2402.03216) disertai sitasi section
- 2d: eksplorasi mandiri satu model “Variation on SBERT” lain di luar slide kuliah
Deadline belum disebutkan eksplisit di brief soal (cuma tanggal rilis 5 Oktober 2026) — 2026-10-12 di frontmatter masih perkiraan (+1 minggu, pola sama seperti Lab Work sebelumnya). Cek SIX/Edunex atau tanya dosen buat tanggal pastinya, lalu update field
deadlinedi sini.
Progress & Catatan Pengerjaan
1. BGE-M3 Hybrid Retrieval
1.1 Prediction
- D1 lebih relevan — “best notebook computer for software developers” cocok secara makna (notebook = laptop, software developer = orang yang nge-program), meski cuma “for” yang sama persis katanya dengan query — D2 “laptop sleeve and carrying case for programming books” share 3 kata persis (laptop, for, programming), tapi isinya sarung/tas laptop + buku programming — bukan laptopnya sendiri
- Sparse retrieval yang paling mungkin salah pilih dokumen — sparse cuma cocokin kata persis tanpa peduli makna, jadi D2 bakal kelihatan “sangat relevan” di matanya padahal nggak
1.2 Dense Score
- D1: ; →
- D2: ; →
1.3 Sparse Score
- D1: hanya token “for” yang muncul di kedua sisi (laptop & programming “tidak muncul” di D1) —
- D2: ketiga token muncul —
1.4 Multi-vector Score
- D1: ambil nilai maksimum tiap baris — laptop→notebook (), for→for (), programming→software () —
- D2 (nilai sudah diberikan di soal): laptop=0.90, for=0.95, programming=0.85 —
1.5 Join
- D1 ranks first () — sesuai prediksi awal, kombinasi hybrid berhasil “menyelamatkan” ranking yang benar meski sparse sendirian salah
1.6 Analysis
- Dense only → D1 menang (0.999 vs 0.769, benar)
- Sparse only → D2 menang (0.3375 vs 0.002, SALAH — tertipu lexical overlap)
- Multi-vector only → D1 menang (2.82 vs 2.70, benar tapi tipis)
- Cari (dengan ) supaya : — —
- Begitu dinaikkan ke ≈1.04 atau lebih (sama/lebih besar dari bobot dense/multi-vec), D2 mulai menang — artinya memilih bobot yang tepat untuk hybrid retrieval itu krusial: kalau sinyal lexical (sparse) dikasih bobot kebesaran, dia bisa membalikkan ranking yang sebenarnya benar dari dense & multi-vec, cuma gara-gara kecocokan kata permukaan — ini sejalan dengan paper BGE-M3: di awal training, matriks bobot sparse () masih diinisialisasi acak sehingga akurasinya rendah/noisy, makanya sengaja dikasih bobot kecil () relatif ke dense & multi-vec () — lihat jawaban 2c buat kutipan persisnya
- Contoh lain query-dokumen sparse > dense: query “error code 0x80070057” (pesan error instalasi Windows)
— dokumen yang memuat string persis
0x80070057gampang ditemukan sparse lewat exact token match; dense bisa gagal karena token alfanumerik kayak gitu jarang muncul di data pretraining, representasinya jadi nggak informatif meski dokumennya justru paling relevan
2. BGE-M3 Paper Exploration
a. Data Curation (Section 3.1 “Data Curation”, Appendix A.1 “Collected Data”)
- Total: 1.2 miliar text pair, 194 bahasa, 2655 kombinasi cross-lingual
- Sumber data unsupervised (≥3): Wikipedia, S2ORC (korpus paper saintifik), xP3 (dataset instruksi multilingual), mC4 (Common Crawl multilingual), CC-News
- Data sintetis GPT-3.5 dibuat untuk long document retrieval task (dataset “MultiLongDoc”) — kutipan: “we generate synthetic data to mitigate the shortage of long document retrieval tasks” — cara buatnya: ambil artikel panjang dari Wikipedia/Wudao/mC4, pilih paragraf acak, GPT-3.5 diminta bikin pertanyaan dari paragraf itu (jadi pasangan question–long-document) — data real-world nggak cukup karena pasangan QA/retrieval berlabel buat dokumen panjang lintas banyak bahasa itu langka — kebanyakan dataset QA publik berbasis paragraf/dokumen pendek
b. Efficient Batching (Section 3.4 “Efficient Batching”)
- Length-grouped sampling: sample satu mini-batch diambil dari grup panjang sequence yang mirip — kutipan: “training instances are sampled from the same group. Due to the similar sequence lengths, it significantly reduces sequence padding.”
— kalau sequence dalam satu batch panjangnya beda jauh, token
[PAD]yang terbuang jadi banyak (komputasi percuma); dengan mengelompokkan sequence sepanjang-mirip, padding yang terbuang berkurang → GPU lebih efisien - Split-batch dengan gradient checkpointing: mini-batch besar dipecah jadi sub-batch kecil, diproses berurutan, pakai gradient checkpointing (buang activation intermediate yang nggak perlu disimpan) buat hemat memori GPU, lalu hasil embedding-nya digabung lagi — buat input 8.192 token, teknik ini naikin batch size lebih dari 20x (Table 10: dari 6 jadi 130 per device)
c. Self-Knowledge Distillation (Section 3.3 “Self-Knowledge Distillation”)
- Bobot di awal/selama fine-tuning: (dense), (sparse), (multi-vec)
- Kenapa nggak disamakan? Kutipan persis: “The random initialization of led to poor accuracy and high at the beginning of the training. In order to reduce the impact of this, we set …” — karena matriks bobot sparse () diinisialisasi acak, skor sparse di awal training masih buruk/noisy (loss-nya tinggi); kalau dibobot sama besar dengan dense, sinyal buruk ini bisa merusak target gabungan buat self-distillation, makanya sengaja diperkecil () selama kualitasnya belum teruji
d. Independent Exploration: INSTRUCTOR (Su et al., 2022, arXiv:2212.09741)
- INSTRUCTOR: satu model embedding yang menerima instruksi bahasa natural sekaligus teks input, menghasilkan embedding yang disesuaikan instruksi tersebut — dilatih multitask learning atas 330 task berbeda sekaligus
- Beda dari SBERT dasar yang cuma di-fine-tune pada satu pasang dataset tetap (NLI lalu STS) buat hasilin satu embedding umum per kalimat — INSTRUCTOR bisa hasilin embedding berbeda-beda untuk kalimat yang sama tergantung instruksi/use-case yang diberikan, tanpa perlu fine-tuning ulang per task baru — diklaim unggul di 70 task evaluasi, dengan parameter jauh lebih sedikit (satu orde besaran lebih kecil) dibanding model terbaik sebelumnya
- Link: https://arxiv.org/abs/2212.09741
Only Answer
1.1
- D1 lebih relevan — D1 memang laptop buat programming (cocok secara makna), D2 cuma sama katanya doang (isinya sarung laptop & buku programming, bukan laptopnya)
- Sparse retrieval yang paling mungkin salah pilih — sparse cuma cocokin kata persis, nggak ngerti makna, jadi gampang ketipu kata sama di konteks beda
1.2 Dense Score
- D1 = 0.999
- D2 = 0.769
1.3 Sparse Score
- D1 = 0.002
- D2 = 0.3375
1.4 Multi-vector Score
- D1 = 2.82
- D2 = 2.70
1.5 Join
- D1 = 3.820, D2 = 3.570 → D1 ranking pertama
1.6 Analysis
- Dense only → D1 menang (benar)
- Sparse only → D2 menang (salah)
- Multi-vector only → D1 menang (benar)
- Yang tertipu lexical overlap: sparse retrieval
- baru D2 mulai ungguli D1 — artinya bobot sparse harus kecil, kalau dibesarin bisa ngerusak ranking yang udah bener dari dense & multi-vec
- Contoh lain sparse > dense: query “error code 0x80070057” — sparse bisa exact-match string error-nya, dense bisa gagal karena token alfanumerik kayak gitu jarang muncul pas training jadi representasinya nggak bagus
2a Data Curation (Section 3.1)
- 1.2 miliar text pair, 194 bahasa
- Sumber: Wikipedia, S2ORC, xP3, mC4, CC-News
- Data sintetis GPT-3.5 buat long document retrieval (dataset “MultiLongDoc”) — dibuat karena pasangan QA/retrieval berlabel buat dokumen panjang lintas banyak bahasa itu langka di data asli
2b Efficient Batching (Section 3.4)
- Length-grouped sampling: kelompokkan sample yang panjangnya mirip ke satu batch — biar token padding yang kebuang dikit, GPU jadi lebih efisien
- Split-batch + gradient checkpointing: batch besar dipecah jadi sub-batch kecil, diproses gantian, buang activation yang nggak perlu disimpan — buat input 8192 token, batch size naik lebih dari 20x (dari 6 jadi 130 per device)
2c Self-Knowledge Distillation (Section 3.3)
- (dense), (sparse) — karena matriks bobot sparse () diinisialisasi acak, skor sparse di awal training masih buruk/noisy, jadi bobotnya dikecilin biar nggak ngerusak target gabungan selama kualitasnya belum teruji
2d Independent Exploration
- Model yang dipilih: INSTRUCTOR (Su et al. 2022, arXiv:2212.09741) — model tunggal yang nerima instruksi bahasa natural + teks, dilatih multitask di 330 task sekaligus; beda dari SBERT yang cuma di-fine-tune NLI→STS buat hasilin satu embedding umum, INSTRUCTOR bisa hasilin embedding beda-beda tergantung instruksinya tanpa perlu fine-tune ulang per task
- Link: https://arxiv.org/abs/2212.09741
Files Terkait
- Brief asli:
_Tugas/1791183331642_Exercise-for-Sentence-Comparison.pdf - Catatan materi terkait: 7_Sentence_Comparison