Residual Stream Calculus
We start with the attention-only path expansion as beautifully explicated in mathematical framework of transformers and extend it to a conditional calculus for nonlinear MLPs, SwiGLU, and mixture-of-experts transformers.
I’m Ji Hun Wang, an Applied Scientist at Amazon working across research and engineering. My interests include post-training, interpretability, and AI safety and alignment. I’m also interested in formal accounts of natural language and the broader relationship between linguistic structure and computation.
Previously, I studied Computer Science and Linguistics at Stanford, completing B.A.S. and M.S. degrees.
We start with the attention-only path expansion as beautifully explicated in mathematical framework of transformers and extend it to a conditional calculus for nonlinear MLPs, SwiGLU, and mixture-of-experts transformers.
The optimization story behind aligning language models with human preferences
About my favorite generative model of all time!
When I prompt an AI model, there are times I wonder how much of a prior I am projecting onto the prompt itself. Even when the task is exploratory by nature or I am not entirely sure how to proceed, I still have some intuitions or guesses as to what might work or what I think is worth trying...
When I prompt an AI model, there are times I wonder how much of a prior I am projecting onto the prompt itself. Even when the task is exploratory by nature or I am not entirely sure how to proceed, I still have some intuitions or guesses as to what might work or what I think is worth trying. In such cases, I try my best to keep my prompt as open-ended as possible, giving as little guidance or “nudge” as possible to steer the model based on my intuition. I do this so that I can tap a greater level of creativity from an AI model that has learned from the whole internet. But regardless of how deliberately direction-agnostic or purportedly unbiased I strive to be, how successful am I?