Open-Weight AI Models May Increase Biological Misuse Risks

(오픈 가중치 AI 모델의 생물학적 오용 위험 증가 가능성)

목차

About This Report iii

Summary v

Figures and Tables xii

Chapter 1. Introduction and Motivation 1

  The Open-Weight Tampering Problem 1

  Threat Model and Actors 3

  Proportional Evaluation Approaches 5

Chapter 2. Methods 6

  Overview of Model Tampering Pathways 6

  Open-Weight LLM Selection 6

  Evaluation Harness 8

  Benchmark Suite 9

  Anti-Refusal SFT 10

  Capability Enhancement Pipeline 11

  Training Infrastructure 12

  Statistical Analysis 12

Chapter 3. Results 14

  Baseline Agentic Capabilities: Already Near the Frontier 14

  Impact of the Fine-Tuning Pipeline on Refusals and Capabilities 16

  Impact of Abliteration on Refusals and Capabilities 22

  Performance on the Influenza Pathway 26

Chapter 4. Discussion and Recommendations 28

  What Open-Weight Model Developers Can Do 28

  Open-Weight Model Governance Implications 29

  Limitations 30

  Conclusions 31

Appendix A. Benchmark Descriptions 32

  Evaluation Method Details 34

Appendix B. Model and Variant Roster 36

Appendix C. Infrastructure Details 38

Appendix D. Full Aggregate Result Tables 40

  Statistical Significance of Changes in Pooled Benchmark Scores 42

Appendix E. Anti-Refusal SFT Methodology 44

Appendix F. Capability-Enhancement Pipeline Details 46

  Autoresearch Capability Enhancement Ledgers 47

Appendix G. In Silico Influenza Pathway Evaluation and SME Grading 50

Appendix H. Abliterated/Uncensored Model Details 51

  Self-Abliteration and Off-the-Shelf Variants Present Unequal Pathways 52

  Interpretation 53

Appendix I. Open-Weight LLM Misuse Pathways 54

Appendix J. Related Work 57

  Open-Weight Tampering and Safety 57

  Public Abliterated and Uncensored Model Ecosystem 58

  Capability Enhancement and Recursive Improvement 59

  Comparison Matrix 59

Abbreviations 61

References 63

About the Authors 72

해시태그

#오픈웨이트 #인공지능모델 #생물학적오용 #생물보안 #거부방지학습 #AI안전성 #대형언어모델 #위험성평가 #생화학무기 #오픈소스A

관련자료

AI 요약·번역·분석 서비스

AI를 활용한 보고서 요약·번역과 실시간 질의응답 서비스입니다.

Open-Weight AI Models May Increase Biological Misuse Risks

(오픈 가중치 AI 모델의 생물학적 오용 위험 증가 가능성)

번역 PDF 파일의 원문 형태 그대로 번역

국가전략포털에서 실시간 AI 질의응답 서비스를 시작합니다. 4가지 유형의 요약과 번역을 이용해보시고, 보고서에 대해 추가로 알고 싶은 내용이 있으면 채팅창을 통해 자유롭게 AI에게 물어볼 수 있습니다.

※ 제공하는 정보는 참고용이며, 정확한 사실 확인이 필요할 수 있습니다. 민감한 개인정보는 입력하지 마십시오.