Extracting Steering Vectors from J space (darshanmakwana412.github.io)

🤖 AI Summary
Recent experiments have demonstrated the potential to extract steering vectors from the Jacobian (J) space of a 1.7B parameter language model (LLM), specifically Qwen3-1.7B. By leveraging the J space, researchers aimed to verbalize the model's intermediate activations and derive general activation vectors that can guide the model's responses to specific concepts using related tokens. The initial findings reveal that this approach can effectively steer the model to exhibit simpler behaviors, such as generating responses in all caps, but it struggles with more complex steering tasks, like generating refusals, which tend to lead to hallucinations and inconsistent outputs. The experiments utilized a technique where concept tokens related to desired behaviors were analyzed to create activation vectors, subsequently averaged to enhance their efficacy. However, steering toward complex behaviors such as refusal showcased significant limitations; the noise and brittleness of the derived activation vectors resulted in outputs that did not consistently reflect the intended steering behavior. This research highlights the advantages of using J space in LLM steering while underscoring the inherent challenges when attempting to articulate and represent complex behaviors through limited token representations.
Loading comments...
loading comments...