A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap,...
Arun Josephraj Arokiaraj, Ze-Kun Wu, A. Koshiyama· 0 citations
The model's hidden state provides a way to read, steer, and check tool choice before a call is made, suggesting that the model's hidden state provides a way to read, steer, and check tool choice before a call is made.
Ze-Kun Wu, Ze-Kun Wang, Seonglae Cho et al.· arXiv.org· 11 citations· ⚡2
This work introduces OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth.
Seonglae Cho, A. Koshiyama· arXiv.org· 0 citations
This work analyzesparse autoencoder features across six models and three SAE families and zero-ablate at full layer depth, finding cross-family claims are sensitive to training methodology, not just activation function or scale.
Seonglae Cho, Zekun Wu, Kleyton Da Costa et al.· arXiv.org· 1 citation
Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
Seonglae Cho, F. Fernandez, Umar Mohammed et al.· 2 citations
This work tracks quantization across 16 models from 8 families under round-to-nearest, seven under AWQ, two under GPTQ and one under GGUF, at 8 down to 2 bits, and measures the margin, the picked option's score minus its best alternative's, which removes the protection a large margin affords.
Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks, so a stronger agent can overturn the result, and the rerun noise bands and the full measurement protocol are reported.
Jiaming Wei, Zekun Wu, A. Koshiyama et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.