Faculty Sponsor: Antonio Laverghetta
Live Poster Session: https://wesleyan.zoom.us/j/94911363126

Ennio Geniblazo
My name is Ennio Geniblazo, and I’m a rising sophomore majoring in Data Science with a minor in Economic. At Wesleyan, I’m researching how LLM’s compute creativity scores through SHAP with Professor Antonio Laverghetta, leveraging a skillset of computer science. Beyond this, I’ve self-studied basic Finance, helping a local cafĂ© in my hometown create its first P&L and business plan geared towards quick optimization. In my free time, I enjoy recreational tennis and exploring the food cart pods in my hometown of Portland, Oregon!
Background
Automated creativity-scoring models can assign originality scores to open-ended responses, but the linguistic features driving these predictions remain difficult to interpret. Furthermore, in creativity studies, expert opinions are expensive. If the industry is to introduce an AI component to cut costs, understanding the process is crucial for maintaining accountability and transparency. This study uses SHapley Additive exPlanations (SHAP) to examine whether token position and spelling classification are associated with token-level contributions to originality predictions.
Methods
Token-level SHAP values were analyzed across four creativity assessment frameworks: AIDE, MAoSS, SCTT, and CPS. For the position analysis, signed and absolute SHAP values were averaged by each token’s sequential position within a response to determine whether particular response sections systematically increased, decreased, or more strongly influenced predicted originality. For the spelling analysis, responses were separated into whitespace-delimited words and classified as correctly spelled or misspelled. These classifications were aligned with the subword tokens produced by each model’s Hugging Face tokenizer. Subword tokens receiving conflicting spelling classifications were excluded, and SHAP values were compared across the remaining spelling categories.
Discussion
The token-position analysis evaluates whether creativity-scoring models rely disproportionately on the beginning, middle, or end of responses. However, later positions contain fewer observations because responses vary in length, requiring cautious interpretation of positional trends. The spelling analysis assesses whether tokens derived from misspelled words contribute differently to originality predictions. Any observed differences may reflect spelling sensitivity, token rarity, or tokenizer behavior rather than a direct model understanding of spelling correctness. Proper nouns, invented words, and technical terms may also be classified as misspellings. Overall, these analyses identify potential positional and linguistic biases in automated creativity scoring and support the development of more transparent and reliable originality models.
egeniblazo_project26