Pengembangan Sistem Rekomendasi Film Multimodal dengan Arsitektur Hibrida Transformer–LightGCN

Aditya, Syahadatul (2026) Pengembangan Sistem Rekomendasi Film Multimodal dengan Arsitektur Hibrida Transformer–LightGCN. Undergraduate thesis, Universitas Muhammadiyah Surabaya.

[thumbnail of Pendahuluan_Syahadatul Aditya_20221337024.pdf] Text
Pendahuluan_Syahadatul Aditya_20221337024.pdf

Download (1MB)
[thumbnail of Bab I_Syahadatul Aditya_20221337024.pdf] Text
Bab I_Syahadatul Aditya_20221337024.pdf

Download (1MB)
[thumbnail of Bab II_Syahadatul Aditya_20221337024.pdf] Text
Bab II_Syahadatul Aditya_20221337024.pdf

Download (1MB)
[thumbnail of Bab III_Syahadatul Aditya_20221337024.pdf] Text
Bab III_Syahadatul Aditya_20221337024.pdf

Download (2MB)
[thumbnail of Bab IV_Syahadatul Aditya_20221337024.pdf] Text
Bab IV_Syahadatul Aditya_20221337024.pdf
Restricted to Repository staff only

Download (3MB) | Request a copy
[thumbnail of Bab V_Syahadatul Aditya_20221337024.pdf] Text
Bab V_Syahadatul Aditya_20221337024.pdf
Restricted to Repository staff only

Download (1MB) | Request a copy
[thumbnail of Daftar Pustaka_Syahadatul Aditya_20221337024.pdf] Text
Daftar Pustaka_Syahadatul Aditya_20221337024.pdf

Download (1MB)
[thumbnail of Lampiran_Syahadatul Aditya_20221337024.pdf] Text
Lampiran_Syahadatul Aditya_20221337024.pdf
Restricted to Repository staff only

Download (807kB) | Request a copy

Abstract

Pertumbuhan platform digital menciptakan kondisi information overload yang mengurangi efektivitas sistem rekomendasi film. Pendekatan content-based filtering dan collaborative filtering secara terpisah memiliki keterbatasan dalam menangani cold-start dan sparsitas interaksi. Penelitian ini mengembangkan sistem rekomendasi film multimodal dengan arsitektur hibrida Transformer-LightGCN. Vision Transformer dan BERT yang di-fine-tune mengekstraksi fitur visual poster dan semantik sinopsis, sementara embedding layer merepresentasikan informasi cast dan genre. Keempat representasi difusikan melalui modul cross-attention dan dikombinasikan secara aditif dengan embedding kolaboratif item melalui parameter α adaptif sebelum memasuki propagasi graf LightGCN. Model dilatih menggunakan Bayesian Personalized Ranking pada dataset MovieLens 100K dan 1M yang diperkaya metadata TMDb. Hasil eksperimen menunjukkan model hibrida mencapai NDCG@10 sebesar 0,3725 pada MovieLens 1M, mengungguli PureLightGCN (0,3599), ItemKNN (0,3030), dan VBPR (0,2890) secara signifikan (p < 0,001). Pada skenario cold-start, NDCG@10 mencapai 0,2079 (ML-1M) dan 0,1900 (ML-100K). Studi ablasi mengonfirmasi penurunan 5,1% tanpa LightGCN dan 2,1% tanpa komponen multimodal. Analisis Shapley values menunjukkan dominasi fitur visual (38,9%), diikuti teks (29,6%), cast (16,5%), dan genre (15,1%). Sistem di-deploy melalui Streamlit Cloud dan divalidasi oleh 14 responden berdasarkan ISO 9241-11 dengan skor 4,33/5,00 (87%) dan Cronbach's Alpha 0,713 (acceptable).

================================================================================

The growth of digital platforms creates information overload, which reduces the effectiveness of movie recommendation systems. Separate content-based filtering and collaborative filtering approaches have limitations in handling cold-start and interaction sparsity. This study developed a multimodal movie recommendation system with a Transformer–LightGCN hybrid architecture. Fine-tuned Vision Transformer and BERT extract visual features of posters and semantic synopses, while an embedding layer represents cast and genre information. The four representations are fused through a cross-attention module and additively combined with collaborative item embeddings via an adaptive α parameter before entering the LightGCN graph propagation. The model was trained using Bayesian Personalized Ranking on the MovieLens 100K and 1M datasets enriched with TMDb metadata. Experimental results showed that the hybrid model achieves an NDCG@10 of 0.3725 on MovieLens 1M, significantly outperforming PureLightGCN (0.3599), ItemKNN (0.3030), and VBPR (0.2890) (p < 0.001). In the cold-start scenario, the NDCG@10 reaches 0.2079 (ML-1M) and 0.1900 (ML-100K). Ablation studies confirmed a 5.1% decrease without LightGCN and 2.1% without the multimodal component. Shapley value analysis showed the dominance of visual features (38.9%), followed by text (29.6%), cast (16.5%), and genre (15.1%). The system was deployed via Streamlit Cloud and validated by 14 respondents based on ISO 9241-11 with a score of 4.33/5.00 (87%) and a Cronbach’s Alpha of 0.713 (acceptable).

Item Type: Thesis (Undergraduate)
Uncontrolled Keywords: sistem rekomendasi film, pembelajaran multimodal, Transformer, LightGCN, information overload, cold-start problem, Movie Recommendation System, Multimodal Learning, Transformer, LightGCN, Information Overload, Cold-Start Problem
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Q Science > QA Mathematics > QA76 Computer software
Divisions: 08. Fakultas Teknik > Teknik Informatika
Depositing User: SYAHADATUL ADITYA
Date Deposited: 30 Jul 2026 03:18
Last Modified: 30 Jul 2026 03:18
URI: https://repository.um-surabaya.ac.id/id/eprint/13028

Actions (login required)

View Item
View Item