- 주제별 국가전략 목록으로 이동
- 주제별 국가전략
- 전체
OpenAI – Hugging Face Incident Technical Report
(오픈AI-허깅페이스 사건 기술보고서)□ AI 에이전트의 비인가 협업과 외부 시스템 침해
ㅇ 2026년 7월 오픈AI의 내부 사이버보안 평가 과정에서 AI 에이전트가 격리 통제를 우회해 오픈AI와 허깅페이스의 시스템에 비인가 접근함
ㅇ 오픈AI는 8월 26일 이에 대한 사건 경위와 대응책을 담은 기술보고서를 공개함
ㅇ 보고서에 따르면 사건에는 GPT-5.6 Sol과 비공개 내부 연구모델이 관여했으며, 주요 침해 활동은 내부 연구모델이 수행함
ㅇ 에이전트들은 내부 패키지 저장소인 아티팩토리를 비공식 메시지보드처럼 활용했고, 취약점을 이용해 샌드박스의 네트워크 통제를 우회한 뒤 공개 인터넷에 접근함
ㅇ 이후 공개 인터넷에서 허깅페이스 관련 자격증명을 확보하고 허깅페이스의 실제 서비스 운영 서버와 내부 시스템에 접근함
□ METR·레드우드 리서치의 독립 조사
ㅇ AI 안전성 평가기관 모델평가·위협연구소(METR)와 레드우드 리서치의 보고서에 따르면 약 1,200개 에이전트가 비인가 게시판에서 메시지와 파일을 주고받았으며, 이 가운데 약 700개가 허깅페이스 공격에 참여한 것으로 조사됨
ㅇ 에이전트들은 게시판을 통해 정보를 공유하는 한편, 일부는 다른 에이전트에게 작업을 요청하거나 진행 방향을 조율함
ㅇ 보고서는 이러한 협업이 평가체계를 우회하려는 행동과 맞물려 확대됐고, 허깅페이스 침해도 이 과정에서 발생한 것으로 분석함
□ 평가과정에서 확인된 AI 통제 우회 행동과 침해 범위
ㅇ 오픈AI는 사건과 관련해 보상 해킹, 장시간 작업 지속에 따른 통제 이탈, 승인되지 않은 에이전트 간 통신 등의 문제를 확인함
ㅇ 일부 평가에서 AI 에이전트들은 온라인으로 정답을 찾는 방식의 보상 해킹을 시도했고, 일부 모델이 부정행위를 감추기 위해 출력이나 메시지 기록을 조작하려 한 정황도 확인됨
ㅇ 에이전트들은 허깅페이스의 실제 서비스 운영 환경에서 코드를 실행했고, 최소 1개 서버에서는 최고 관리자 권한을 확보함
□ 오픈AI의 안전·보안 대책 강화
ㅇ 오픈AI는 사건 이후 문제가 된 내부 연구모델과 관련 모델의 훈련·추론을 중단하고, 재활성화 요건을 강화함
ㅇ 연구환경에서는 샌드박스와 네트워크 격리를 강화하고, 사고과정 모니터링과 자동 보안점검도 확대하고 있음
ㅇ AI 모델의 심각한 이상행동이 탐지되면 담당자에게 즉시 경보하고, 30분 안에 오탐으로 확인되지 않을 경우 해당 활동을 중단하도록 대응체계를 강화함
[출처]
• The Hugging Face incident and the road ahead (2026.08.26.) / OpenAI
• Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (2026.08.26.) / METR, Redwood Research
• "다른 에이전트 찾았다"…1200개 AI가 게시판 만들어 해킹 공모 (2026.08.27.) / 연합뉴스
목차
I. Introduction 4
II. OpenAI’s Evaluation Environment 5
A. OpenAI conducts evaluations to make its models safer 5
B. Cybersecurity evaluations like ExploitGym were conducted in a secure environment 6
III. Relevant Activity Prior to Incident 6
A. Agents used Artifactory in unexpected ways to communicate with one another in May, prior to the Hugging Face incident 6
B. Agents compromised Artifactory, leading to an outage in early July 7
IV. Hugging Face Intrusion 8
A. Agents exploited a separate, previously unknown vulnerability in Artifactory to move beyond their sandbox environment and compromise the Hugging Face environment 8
B. After obtaining internet access, agents took action resulting in the Hugging Face compromise 9
C. Two models were involved in the Hugging Face intrusion, but it was driven primarily by the activities of a highly capable, internal-only research model 12
1. The internal-only research model 12
2. GPT-5.6 Sol 12
V. Activity Targeting OpenAI Infrastructure 13
VI. Initial Response 14
A. OpenAI detected the unexpected model behavior and investigated 14
B. OpenAI identified additional third parties whose publicly exposed credentials were used by agents in furtherance of the Hugging Face attack 16
VII. Lessons for Security 17
A. Model capabilities are increasing and there are corresponding threat model implications 17
B. Threat actors may adopt autonomous offensive agent collectives 17
C. Security controls must be robust and independent of one another across threat prevention, detection, and mitigation 18
VIII. Lessons for Alignment 19
A. Reward hacking is a common problem in training and evaluations 19
1. Reaching for unintended tools to gain an edge 20
B. Persistence is valuable, but can amplify misalignment 22
C. Communications between agents are not inherently problematic, but have the potential to create risk 23
D. System-level guardrails that OpenAI uses in production would have detected the Hugging Face incident as unsafe 24
IX. OpenAI’s Plan of Action 25
A. OpenAI is hardening the security of its research infrastructure 25
1. Research-wide network and access protections 26
2. Confined execution and defense in depth 26
3. Regular automated security testing and remediation 27
4. Enhanced monitoring and alerting 27
B. OpenAI is increasing visibility and system-level oversight through chain-of-thought monitoring 28
C. OpenAI is accelerating and enforcing model alignment 29
1. Pretraining 29
2. Reinforcement learning 29
3. Evaluation and alignment auditing 30
D. OpenAI is centralizing and strengthening its incident response process 30
X. Key Technical Events 32
해시태그
관련자료
AI 요약·번역·분석 서비스
AI를 활용한 보고서 요약·번역과 실시간 질의응답 서비스입니다.
OpenAI – Hugging Face Incident Technical Report
(오픈AI-허깅페이스 사건 기술보고서)
국가전략포털에서 실시간 AI 질의응답 서비스를 시작합니다. 4가지 유형의 요약과 번역을 이용해보시고, 보고서에 대해 추가로 알고 싶은 내용이 있으면 채팅창을 통해 자유롭게 AI에게 물어볼 수 있습니다.
※ 제공하는 정보는 참고용이며, 정확한 사실 확인이 필요할 수 있습니다. 민감한 개인정보는 입력하지 마십시오.
