Skip to main content
Seminar | Computing, Environment and Life Sciences

Resource-friendly alignment in language models: from reward modeling to preference learning

Trillion Parameter Consortium (TPC) Seminar

Abstract: Human preference alignment has become indispensable in large language model (LLM) post-training schemes. Despite its significance, preparing adequate reward models and large amounts of high-quality data is a crucial and resource-intensive process that must precede alignment-tuning. This is especially challenging for tasks or languages with limited data accessibility, making large-scale alignment-tuning difficult. 

This talk reviews resource-efficient, end-to-end schemes for aligning LLMs in language-specific or task-specific contexts. In particular, it covers the two main stages of alignment-tuning: reward modeling and preference learning. For reward modeling, we demonstrate the versatility of English-based reward models across different languages, highlighting the conditions that enable them to function as language-agnostic models. We then explore the mechanism of Odds-Ratio Preference Optimization (ORPO) for either directly aligning pre-trained LLMs or existing instruction-following LLMs without requiring vast amounts of supervised fine-tuning data, and present real-world, task-specific use cases of ORPO.

BIO: Jiwoo Hong is an MSc student at KAIST AI under the supervision of Professor James Thorne. He completed his undergraduate degree at SKKU, majoring in Statistics and Industrial Engineering.