{"id":1181465,"date":"2026-08-12T02:07:08","date_gmt":"2026-08-12T09:07:08","guid":{"rendered":"https:\/\/research.codeghost.online\/en-us\/research\/?post_type=msr-research-item&#038;p=1181465"},"modified":"2026-08-12T02:12:02","modified_gmt":"2026-08-12T09:12:02","slug":"exposing-weak-links-in-multi-agent-systems-under-adversarial-prompting","status":"publish","type":"msr-research-item","link":"https:\/\/research.codeghost.online\/en-us\/research\/publication\/exposing-weak-links-in-multi-agent-systems-under-adversarial-prompting\/","title":{"rendered":"Exposing Weak Links in Multi-Agent Systems under Adversarial Prompting"},"content":{"rendered":"\n\n\n<p class=\"wp-block-paragraph\">LLM-based agents are increasingly deployed in multi-agent systems (MAS). As these systems move toward real-world applications, their security becomes paramount. Existing research largely evaluates single-agent security, leaving a critical gap in understanding the vulnerabilities introduced by multi-agent design. However, existing evaluation approaches fall short due to lack of unified frameworks and metrics focusing on unique rejection modes in MAS. We present SafeAgents, a unified and extensible framework for fine-grained security assessment of MAS. SafeAgents systematically exposes how design choices such as plan construction strategies, inter-agent context sharing, and fallback behaviors affect susceptibility to adversarial prompting. We introduce Dharma, a diagnostic measure that helps identify weak links within multi-agent pipelines. Using SafeAgents, we conduct a comprehensive study across five widely adopted multi-agent architectures (centralized, decentralized, and hybrid variants) on four datasets spanning web tasks, tool use, and code generation. Our central finding is that MAS security failures are not random&#8212;they are predictable consequences of specific, identifiable design choices. We trace failures to three recurring weak links: atomic-instruction delegation that hides harmful intent from sub-agents, missing planner fallbacks that turn refusals into execution, and stratified plans executed without re-evaluation. These results argue for security-aware design rather than post-hoc safeguards. Link to code is https:\/\/github.com\/microsoft\/SafeAgents<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>LLM-based agents are increasingly deployed in multi-agent systems (MAS). As these systems move toward real-world applications, their security becomes paramount. Existing research largely evaluates single-agent security, leaving a critical gap in understanding the vulnerabilities introduced by multi-agent design. However, existing evaluation approaches fall short due to lack of unified frameworks and metrics focusing on unique [&hellip;]<\/p>\n","protected":false},"featured_media":0,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","msr-author-ordering":[{"type":"text","value":"Nirmit Arora","user_id":0},{"type":"text","value":"Sathvik Joel","user_id":0},{"type":"text","value":"Ishan Kavathekar","user_id":0},{"type":"text","value":"Palak LNU","user_id":0},{"type":"user_nicename","value":"Rohan Gandhi","user_id":"42372"},{"type":"user_nicename","value":"Yash Pandya","user_id":"44036"},{"type":"user_nicename","value":"Tanuja Ganu","user_id":"38883"},{"type":"text","value":"Aditya Kanade ","user_id":0},{"type":"user_nicename","value":"Akshay Nambi","user_id":"38169"}],"msr_publishername":"","msr_publisher_other":"","msr_booktitle":"","msr_chapter":"","msr_edition":"","msr_editors":"","msr_how_published":"","msr_isbn":"","msr_issue":"","msr_journal":"AAMAS SE","msr_number":"","msr_organization":"","msr_pages_string":"","msr_page_range_start":"","msr_page_range_end":"","msr_series":"","msr_volume":"","msr_copyright":"","msr_conference_name":"","msr_doi":"","msr_arxiv_id":"","msr_mag_id":"","msr_other_authors":"","msr_other_contributors":"","msr_speaker":"","msr_award":"","msr_affiliation":"","msr_institution":"","msr_host":"","msr_version":"","msr_duration":"","msr_release_tracker_id":"","msr_highlight_type":"","msr_date_display_format":"","msr_main_download_label":"","msr_external_link_label":"","msr_doi_label":"","msr_published_date":"2026-05","msr_startdate":"","msr_presentation_date":"","msr_highlight_text":"","msr_notes":"","msr_longbiography":"","msr_publicationurl":"","msr_external_url":"","msr_secondary_video_url":"","msr_conference_url":"","msr_journal_url":"","msr_year":2026,"msr_month":5,"msr_day":0,"msr_microsoftintellectualproperty":false,"msr_pub_id":"","msr_publication_uploader":[{"type":"file","title":"safeagents-aamas26.pdf","label_id":243132,"id":1181466,"viewUrl":"https:\/\/research.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/08\/safeagents-aamas26.pdf"}],"msr_related_uploader":[],"msr_original_fields_of_study":[],"msr_s2_paper_id":"","msr_s2_pdf_url":"","msr_citation_count_updated":"","msr_citation_count":0,"msr_influential_citations":0,"msr_reference_count":0,"msr_s2_open_access":false,"msr_s2_author_ids":[],"msr_pub_ids":[],"msr_hide_image_in_river":0,"footnotes":""},"msr-research-highlight":[],"research-area":[13556,13558],"msr-publication-type":[193715],"msr-publisher":[],"msr-publication-cta":[],"msr-focus-area":[243210,243573],"msr-locale":[268875],"msr-post-option":[],"msr-field-of-study":[270512,267246],"msr-conference":[],"msr-journal":[],"msr-impact-theme":[261676],"msr-pillar":[],"class_list":["post-1181465","msr-research-item","type-msr-research-item","status-publish","hentry","msr-research-area-artificial-intelligence","msr-research-area-security-privacy-cryptography","msr-focus-area-machine-learning","msr-focus-area-trustworthy-computing","msr-locale-en_us","msr-field-of-study-llm-security","msr-field-of-study-multi-agent-systems"],"msr_publishername":"","msr_edition":"","msr_affiliation":"","msr_published_date":"2026-05","msr_host":"","msr_duration":"","msr_version":"","msr_speaker":"","msr_other_contributors":"","msr_booktitle":"","msr_pages_string":"","msr_chapter":"","msr_isbn":"","msr_journal":"AAMAS SE","msr_volume":"","msr_number":"","msr_editors":"","msr_series":"","msr_issue":"","msr_organization":"","msr_how_published":"","msr_notes":"","msr_highlight_text":"","msr_release_tracker_id":"","msr_original_fields_of_study":"","msr_download_urls":"","msr_external_url":"","msr_secondary_video_url":"","msr_longbiography":"","msr_microsoftintellectualproperty":0,"msr_main_download":"","msr_publicationurl":"","msr_doi":"","msr_publication_uploader":[{"type":"file","title":"safeagents-aamas26.pdf","label_id":243132,"id":1181466,"viewUrl":"https:\/\/research.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/08\/safeagents-aamas26.pdf"}],"msr_related_uploader":[],"msr_citation_count":0,"msr_citation_count_updated":"","msr_s2_paper_id":"","msr_influential_citations":0,"msr_reference_count":0,"msr_arxiv_id":"","msr_s2_author_ids":[],"msr_s2_open_access":false,"msr_s2_pdf_url":null,"msr_attachments":[{"id":1181466,"url":"https:\/\/research.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/08\/safeagents-aamas26.pdf"}],"msr-author-ordering":[{"type":"text","value":"Nirmit Arora","user_id":0,"rest_url":false},{"type":"text","value":"Sathvik Joel","user_id":0,"rest_url":false},{"type":"text","value":"Ishan Kavathekar","user_id":0,"rest_url":false},{"type":"text","value":"Palak LNU","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Rohan Gandhi","user_id":42372,"rest_url":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Rohan Gandhi"},{"type":"user_nicename","value":"Yash Pandya","user_id":44036,"rest_url":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Yash Pandya"},{"type":"user_nicename","value":"Tanuja Ganu","user_id":38883,"rest_url":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Tanuja Ganu"},{"type":"text","value":"Aditya Kanade","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Akshay Nambi","user_id":38169,"rest_url":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Akshay Nambi"}],"msr_impact_theme":["Trust"],"msr_research_lab":[],"msr_event":[],"msr_group":[],"msr_project":[],"publication":[],"video":[],"msr-tool":[],"msr_publication_type":"article","related_content":[],"_links":{"self":[{"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1181465","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item"}],"about":[{"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-research-item"}],"version-history":[{"count":5,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1181465\/revisions"}],"predecessor-version":[{"id":1181473,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1181465\/revisions\/1181473"}],"wp:attachment":[{"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1181465"}],"wp:term":[{"taxonomy":"msr-research-highlight","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-highlight?post=1181465"},{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1181465"},{"taxonomy":"msr-publication-type","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-type?post=1181465"},{"taxonomy":"msr-publisher","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publisher?post=1181465"},{"taxonomy":"msr-publication-cta","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-cta?post=1181465"},{"taxonomy":"msr-focus-area","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-focus-area?post=1181465"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1181465"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=1181465"},{"taxonomy":"msr-field-of-study","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-field-of-study?post=1181465"},{"taxonomy":"msr-conference","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-conference?post=1181465"},{"taxonomy":"msr-journal","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-journal?post=1181465"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=1181465"},{"taxonomy":"msr-pillar","embeddable":true,"href":"https:\/\/research.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-pillar?post=1181465"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}