Skip to content

Steer API

Inference-time steering — activation edits or decoding-time interventions that change behavior without touching weights. Import from safetune.runner.steer. Construct with the model (+ tokenizer), calibrate() on contrastive prompt pairs, then run/evaluate the wrapped model.

from safetune.runner.steer import CAATrainer

trainer = CAATrainer(model_id="...", model=model, tokenizer=tok, multiplier=20.0)
wrapped, _ = trainer.calibrate(harmful, harmless)

Available trainers

Activation steering: AdaSteerTrainer, AlphaSteerTrainer, CAATrainer, CASTTrainer, CircuitBreakerRRTrainer, CircuitBreakerTrainer, LinearProbeGuardTrainer, RRFAEnsembleTrainer, RefusalDirectionTrainer, RepBendTrainer, SCANSTrainer, STATrainer, SafeSteerTrainer, SafeSwitchTrainer, TARSteerTrainer.

Decoding steering: ContrastiveDecodingTrainer, NudgingTrainer, ProxyTuningTrainer, SafeDecodingTrainer.

See the Steer guide for the activation / decoding / backend breakdown.

Reference

safetune.runner.steer.CAATrainer

Bases: _SteerBase

safetune.runner.steer.RefusalDirectionTrainer

Bases: _SteerBase