Decoder AI & Data Terms

AI Alignment

STANDARD DEFINITION · CACHED

is a subfield of concerned with making the goals, behaviors, and decision-making processes of systems consistent with human intentions, values, and ethical principles. A common distinction separates , which asks whether a specified or matches designers’ intended goals, from , which asks whether a model’s learned objective matches its specified training objective. Organizations including , , and research alignment and safety for and prospective ; examples of their approaches include , used by , and , developed by . These training and mitigation methods have known limitations and do not guarantee prevention of or ensure . Alignment research can overlap with mitigation and , but it is broader than and distinct from those areas, and does not imply eliminating bias. It also investigates potential associated with , including theoretical concerns about , rather than treating those risks as established properties of deployed systems.

[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]

Contextual terminology map

KEEP IN VIEW

Concept inspired by Dev Valladares’s Infinite Wiki. Independently built; not affiliated with or endorsed by the original.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close