Policy Optimization for RL and LLM Alignment — Theory
The optimization story behind aligning language models with human preferences
I’m Ji Hun Wang, an Applied Scientist at Amazon working across research and engineering. My interests include post-training, interpretability, and AI safety and alignment. I’m also interested in formal accounts of natural language and the broader relationship between linguistic structure and computation.
Previously, I studied Computer Science and Linguistics at Stanford, completing B.A.S. and M.S. degrees.
The optimization story behind aligning language models with human preferences
About my favorite generative model of all time!
When I prompt an AI model, there are times I wonder how much of a prior I am projecting onto the prompt itself. Even when the task is exploratory by nature or I am not entirely sure how to proceed, I still have some intuitions or guesses as to what might work or what I think is worth trying...
When I prompt an AI model, there are times I wonder how much of a prior I am projecting onto the prompt itself. Even when the task is exploratory by nature or I am not entirely sure how to proceed, I still have some intuitions or guesses as to what might work or what I think is worth trying. In such cases, I try my best to keep my prompt as open-ended as possible, giving as little guidance or “nudge” as possible to steer the model based on my intuition. I do this so that I can tap a greater level of creativity from an AI model that has learned from the whole internet. But regardless of how deliberately direction-agnostic or purportedly unbiased I strive to be, how successful am I?