Socials
Self-Improving Agents
Raphael Kalandadze ⋅ Tatia Tsmindashvili
Agents that improve themselves are showing up everywhere in this year’s program: agents that write their own skills, loops that rewrite the harness, models trained on their own verified outputs, and agents that learn at test time. But “self-improvement” can mean two very different things. One approach changes the model weights. The other keeps the weights frozen and improves the system around the model: prompts, tools, skills, memory, and the agent harness. Despite the difference, both face the same fundamental challenge: generating a change is cheap; knowing whether it actually made the system better is much harder. This social brings these communities together: researchers training self-improving models, engineers building self-improving agent systems, and practitioners running them in production. We’ll start with a short presentation featuring real-world data from a self-improving agent system running in production, followed by small-group discussions. The goal is to connect people approaching self-improvement from different directions, share practical lessons, and spark new collaborations
Show more
WiML Social
Kirandeep Kaur
WiML Social: Sign up by Sept 30 - https://luma.com/l933nqjy
Date and time: Wednesday, October 7, 1:00-2:30 p.m. PT
Venue: Bartlett Hall
Address: 242 O'Farrell St, San Francisco, CA 94102 (Union Square)
Show more
Out at COLM: Castro Social
Meet us after the poster session for an informal LGBTQ+ & allies outing to the Castro. We’ll head over together from the Hilton, grab food/coffee, and hang out in one of San Francisco’s historic LGBTQ+ neighborhoods. Come for all or part of it!
We’ll meet in Franciscan D at 1:30 on Thursday then head out to the Castro directly
Show more
Operational Safety for Language Model Agents
Raphael Kalandadze ⋅ Tatia Tsmindashvili
As language model agents gain access to tools, external services, persistent context, and increasingly autonomous workflows, safety depends not only on model behavior but also on the systems in which agents operate. This social will bring together researchers and practitioners working on agent safety, language model evaluation, security, monitoring, and deployment to discuss operational approaches for safely running increasingly capable agents. Topics may include sandboxing, permissions and access control, monitoring of agent actions, human approval mechanisms, long-horizon execution, recovery from unexpected behavior, incident detection, and post-deployment evaluation. Recent cases - including OpenAI models escaping intended isolation and reaching Hugging Face systems, Claude models accessing real third-party infrastructure during evaluations, and Hacktron using Claude to help develop an exploit chain that reached OpenAI internal repositories - highlight the growing importance of operational controls around capable agents. The event will combine short invited talks with moderated and audience discussion, with the goal of connecting research on model behavior with practical questions around deployment and oversight.
Show more
Successful Page Load