Abstract
Political leadership has long been evaluated through social expectations surrounding authority, professionalism, electability, and legitimacy. As large language models (LLMs) become increasingly integrated into political communication, campaigns, journalism, and information search, understanding how these systems evaluate political candidates has become an important scholarly question. Despite growing research on algorithmic bias, comparatively little attention has been devoted to AI-generated assessments of political candidates or whether those assessments vary systematically across candidate identities. This study examines whether commercially available generative AI systems produce systematically different evaluations of otherwise comparable political candidates when sexual orientation, gender identity, and leadership presentation vary while substantive qualifications remain constant. Using a standardized vignette experiment, eighteen candidate profiles were evaluated by six widely used large language models across repeated independent sessions, yielding 540 AI-generated candidate evaluations. The analysis compares assessments of electability, leadership, competence, political strength, and overall candidate quality across candidate conditions.

![Author ORCID: We display the ORCID iD icon alongside authors names on our website to acknowledge that the ORCiD has been authenticated when entered by the user. To view the users ORCiD record click the icon. [opens in a new tab]](https://preprints.apsanet.org/engage/assets/public/apsa/logo/orcid.png)