Resource-friendly alignment in language models: from reward modeling to preference learning
Events section menu
Abstract: Human preference alignment has become indispensable in large language model (LLM) post-training schemes. Despite its significance, preparing adequate reward models and large amounts of high-quality data is a crucial and resource-intensive process that must precede alignment-tuning. This is especially challenging for tasks or languages with limited data accessibility, making large-scale alignment-tuning difficult.
This talk reviews resource-efficient, end-to-end schemes for aligning LLMs in language-specific or task-specific contexts. In particular, it covers the two main stages of alignment-tuning: reward modeling and preference learning. For reward modeling, we demonstrate the versatility of English-based reward models across different languages, highlighting the conditions that enable them to function as language-agnostic models. We then explore the mechanism of Odds-Ratio Preference Optimization (ORPO) for either directly aligning pre-trained LLMs or existing instruction-following LLMs without requiring vast amounts of supervised fine-tuning data, and present real-world, task-specific use cases of ORPO.
BIO: Jiwoo Hong is an MSc student at KAIST AI under the supervision of Professor James Thorne. He completed his undergraduate degree at SKKU, majoring in Statistics and Industrial Engineering.