I fed 195 of the best films ever made to a model and charted where the feeling goes
Ten things the dialogue of the IMDb Top 250 says about how great films handle jokes, tension, endings and grief, each with a chart. Then the method, the caveats, and what did not work.
I had a folder in Google Drive with the English subtitles for the IMDb Top 250 and no plan for it. The idea I kept coming back to is from granular synthesis. You take a sound, cut it into tiny overlapping grains, and the one continuous thing becomes a pile of small things you can measure and rearrange. Do that to a film and you get a few dozen five-minute windows of dialogue, each small enough to ask a simple question about. How tense is this? How funny? Is anybody grieving?
So I did it. 195 films, 10,129 windows, 3.8 million words of dialogue, eight emotion scores per window plus a dominant emotion, all placed on a shared clock that runs from 0% to 100% of the film so a 90-minute comedy and a three-hour epic can be laid on top of each other. The ten things I found come first. The method, the caveats, and the stuff that did not work are after that, for the people who want to argue.
Ten things the charts say
1. Great films front-load the jokes
Humor peaks about a quarter of the way in and sinks below the film's own average by the end. This is dialogue humor only, so a sight gag or a Buster Keaton fall is invisible here. The words still tell one consistent story: you get the laughs early and you pay for them later.
Humor peaks early and never comes back
Mean humor score across all films at each point of narrative time. Each film is z-scored first, so this is shape, not level. The band is one standard deviation.
Table view
| runtime | humor |
|---|---|
| 5% | 0.16 |
| 14% | 0.37 |
| 25% | 0.44 |
| 35% | 0.21 |
| 45% | 0.22 |
| 55% | 0.07 |
| 65% | -0.07 |
| 75% | -0.21 |
| 85% | -0.42 |
| 95% | -0.75 |
| 100% | -0.77 |
2. Tension climbs the whole film and then falls off a cliff
On average, tension rises for 88% of the runtime and then drops harder than anything else in the data. That drop is the resolution. Arousal does the same thing a beat behind. It is the most consistent shape in the whole set.
Tension rises for nearly the entire film, then collapses
Mean tension and arousal across all films, z-scored per film. The last few percent of runtime is where the corpus moves together.
Table view
| runtime | tension | arousal |
|---|---|---|
| 5% | -0.52 | -0.53 |
| 14% | -0.26 | -0.19 |
| 25% | -0.19 | -0.16 |
| 35% | -0.08 | -0.10 |
| 45% | -0.02 | -0.05 |
| 55% | 0.10 | 0.07 |
| 65% | 0.25 | 0.19 |
| 75% | 0.39 | 0.36 |
| 85% | 0.47 | 0.35 |
| 95% | 0.05 | 0.21 |
| 100% | -0.64 | -0.33 |
3. The ending is the one place these films agree
For most of the runtime the films disagree about where their beats go, which is why the average curves are so flat in the middle. Then the last tenth arrives and they line up. Warmth peaks there in 31% of films and valence in 30%. Humor peaks there in 2%.
Where the peak lands: the share of films that save it for the last 10%
For each axis, the percentage of films whose curve reaches its maximum in the final tenth of runtime.
Table view
| films peaking in last 10% | |
|---|---|
| warmth | 31% |
| valence | 30% |
| grief | 23% |
| moral weight | 22% |
| tension | 22% |
| conflict | 17% |
| arousal | 16% |
| humor | 2% |
4. Warmth, grief and moral weight all get heavier as the film goes
On the raw 0 to 4 scale, the average film gains about half a point of warmth, grief and moral weight between its first tenth and its last tenth. Humor loses more than half a point. Valence and conflict barely move on average, which hides a lot of individual movement. That is the next chart.
Opening level versus closing level, by axis
Average score in the first 10% of runtime against the last 10%, on the 0 to 4 rubric scale. Sorted by how much each axis rises.
Table view
| axis | first 10% | last 10% | change |
|---|---|---|---|
| warmth | 1.40 | 1.91 | +0.51 |
| moral weight | 1.90 | 2.37 | +0.47 |
| grief | 1.05 | 1.51 | +0.46 |
| arousal | 2.52 | 2.80 | +0.28 |
| tension | 2.21 | 2.47 | +0.27 |
| valence | 1.48 | 1.50 | +0.01 |
| conflict | 1.83 | 1.83 | +0.00 |
| humor | 1.80 | 1.16 | -0.64 |
5. Fewer than half of them end happier than they start
93 of the 195 end with higher valence than they opened with. 102 end lower. The happy ending is not the default in this canon, and a lot of the films you would call uplifting get there through an ending that is quieter and sadder than their opening.
Ending valence minus opening valence, one bar per half point
Each film's last-10% average minus its first-10% average. Bars to the right of zero are films that end up.
Table view
| change in valence | films |
|---|---|
| -2.5 to -2.0 | 0 |
| -2.0 to -1.5 | 4 |
| -1.5 to -1.0 | 9 |
| -1.0 to -0.5 | 36 |
| -0.5 to 0 | 53 |
| 0 to +0.5 | 43 |
| +0.5 to +1.0 | 30 |
| +1.0 to +1.5 | 15 |
| +1.5 to +2.0 | 2 |
| +2.0 to +2.5 | 3 |
6. There is no secret set of story shapes
I clustered the curves looking for the six or seven shapes everyone hopes to find. The best split on all eight axes has a silhouette score of 0.04. On valence alone it is 0.08. A score of 1 would mean clean separation, and 0 means the clusters are no tighter than a random cut. The honest reading is that these films sit on a continuum. It could be that my axes are the wrong ones, and I say more about that below, but on this data there are tendencies, not types.
How separable the clusters are, for every number of clusters I tried
Mean silhouette from k-means on the z-scored curves. Nothing gets near the range where you would call the groups real.
Table view
| k | all eight axes | valence only |
|---|---|---|
| k = 2 | 0.036 | 0.082 |
| k = 3 | 0.036 | 0.078 |
| k = 4 | 0.033 | 0.069 |
| k = 5 | 0.03 | 0.066 |
| k = 6 | 0.031 | 0.063 |
7. The one split that does show up is "ends up" versus "ends down"
When you force the valence curves into two groups, the only thing the groups disagree about is the last few minutes. 90 films rise at the close. 105 fall. Everything before that looks the same.
The two valence clusters, and the only place they differ
Mean z-scored valence curve of each cluster. Identical for 90% of the runtime, then they part.
Table view
| runtime | ends up (90) | ends down (105) |
|---|---|---|
| 5% | 0.16 | 0.31 |
| 14% | -0.24 | 0.19 |
| 25% | 0.09 | 0.12 |
| 35% | -0.23 | 0.30 |
| 45% | -0.10 | 0.11 |
| 55% | -0.18 | -0.05 |
| 65% | -0.16 | -0.10 |
| 75% | -0.22 | -0.39 |
| 85% | -0.11 | -0.36 |
| 95% | 0.89 | -0.61 |
| 100% | 1.41 | 0.04 |
8. The peak-end rule does not predict which films rank highest
Kahneman's rule says we remember an experience by its most intense moment and its ending. I tested it. Correlate each film's valence peak, its ending, and the average of the two against its place on the list. All three land near zero. The one valence number that moves with rank is the opening, and it moves the wrong way for a feel-good theory: films that open darker rank higher.
Valence measures against IMDb rank
Spearman correlation between each measure and rank, with rank flipped so a positive number means more of this goes with a higher place on the list. Values above 0.14 in either direction clear p below .05, and are marked.
Table view
| rho vs. rank | |
|---|---|
| mean | -0.12 |
| peak | -0.04 |
| end | -0.03 |
| peak end | -0.05 |
| start | -0.16 * |
9. Tension and moral stakes go with a higher rank. Jokes go the other way.
Average tension across the whole film correlates with rank at +0.22, moral weight at +0.17. Humor is negative and not significant. These are modest numbers inside an already elite list, so read them as which way the wind blows, not as a formula. The wind blows toward stakes.
A film's average level on each axis, against its rank
Same test as Chart 8, one bar per axis, using the film's mean score across its whole runtime.
Table view
| rho vs. rank | |
|---|---|
| tension | +0.22 * |
| moral weight | +0.17 * |
| arousal | +0.09 |
| conflict | +0.07 |
| grief | +0.05 |
| warmth | +0.05 |
| humor | -0.10 |
| valence | -0.12 |
10. The rater passed the sniff test
Before you trust any of the above you want to know the model reads dialogue the way a person would. The most positive film on the list is Singin' in the Rain. The funniest is Monty Python and the Holy Grail. The most tense is Dr. Strangelove. Downfall, Full Metal Jacket and Oldboy sit at the bottom of valence, and the loudest film in the set is The Wolf of Wall Street, which anyone who has seen it will accept without a chart.
The top three films on three axes
Highest average score across the whole film. If these looked wrong, nothing above would be worth reading.
Table view
| film | mean level | |
|---|---|---|
| most positive #1 | Singin' in the Rain (1952) | 2.89 |
| most positive #2 | WALL·E (2008) | 2.51 |
| most positive #3 | My Neighbor Totoro (1988) | 2.48 |
| funniest #1 | Monty Python and the Holy Grail (1975) | 3.72 |
| funniest #2 | Some Like It Hot (1959) | 3.16 |
| funniest #3 | Finding Nemo (2003) | 3.09 |
| most tense #1 | Dr. Strangelove or- How I Learned to Stop Worrying and Love the Bomb (1964) | 3.61 |
| most tense #2 | Avengers- Infinity War (2018) | 3.49 |
| most tense #3 | The Dark Knight (2008) | 3.47 |
How I did it
The subtitles came from a folder of 200 SRT files, one per film, ranked at the time they were collected. Five came out before any analysis: four silent films that only have intertitles (Modern Times, City Lights, Metropolis, The Kid) and Kill Bill: The Whole Bloody Affair, whose subtitle file has 230 cues for a 232-minute cut and is clearly incomplete. Kill Bill Vol. 1 stayed. That leaves 195.
Each file was parsed, stripped of formatting tags and the ads subtitle sites inject, and cut into windows.
Five-minute windows, every two and a half minutes
The grain size. Two minutes left too many windows with almost no dialogue during action and montage. Ten minutes gave only a dozen points per film, too few to see a shape.
Table view
| setting | value |
|---|---|
| window length | 5 min |
| hop | 2.5 min |
| overlap | 50% |
| windows per film | about 50 |
Every window's dialogue went to Jev, TypeSafe AI's evaluation model, through the Vercel AI Gateway, where it is free. Jev is not a chat model. You give it a piece of text and a set of typed questions, and it returns a calibrated answer for each without generating any prose. For a score question you supply the rungs of a rubric in order, and it returns a probability for each rung plus an interpolated score. That is a better fit for this job than asking a chat model to emit JSON and hoping the numbers mean the same thing on window 9,000 as they did on window 1. The model saw only the dialogue. No title, no year. If it knew it was reading The Godfather it would rate from memory of the film instead of from the text, and the measurement would be gone.
The eight axes, each a five-rung rubric scored 0 to 4, and the theory each one is anchored to:
| Axis | Anchored to | |---|---| | valence | Russell's circumplex model of affect | | arousal | Russell, and Zillmann's excitation transfer | | tension | Brewer and Lichtenstein's structural affect; Huron's theory of expectation | | conflict | Zillmann's affective disposition theory | | warmth | kama muta, the "being moved" research from Fiske and Menninghaus | | grief | Bonanno's grief trajectories; Herman's stages of recovery | | moral weight | Haidt's moral foundations; Shay's moral injury | | humor | its own axis, because humor is not the opposite of sadness |
Plus a ninth question that picks one dominant emotion from Plutchik's eight, or neutral. The pipeline is Bun and TypeScript, eight calls in parallel, and it finished 10,129 windows in about twenty minutes with 1,690 retries and no failures. The gateway was having a day. I had Claude Code write most of it.
Then the normalization. Every film's windows were placed on its own runtime, resampled to 100 points from 0% to 100%, and z-scored per film per axis. The z-score is the choice that makes the curves comparable. It throws away how tense a film is on average and keeps only where in the film the tension goes. The raw 0 to 4 levels are kept separately for the questions that need them.
All eight average curves
The same view as Charts 1 and 2 for every axis. Flat middles, busy endings.
Table view
| runtime | valence | arousal | tension | conflict | warmth | grief | moral weight | humor |
|---|---|---|---|---|---|---|---|---|
| 5% | 0.25 | -0.53 | -0.52 | -0.29 | -0.43 | -0.32 | -0.50 | 0.16 |
| 14% | 0.04 | -0.19 | -0.26 | 0.02 | -0.22 | -0.17 | -0.17 | 0.37 |
| 25% | 0.13 | -0.16 | -0.19 | -0.05 | -0.11 | -0.15 | -0.23 | 0.44 |
| 35% | 0.04 | -0.10 | -0.08 | 0.01 | -0.14 | -0.11 | -0.07 | 0.21 |
| 45% | 0.03 | -0.05 | -0.02 | 0.01 | 0.08 | -0.06 | -0.03 | 0.22 |
| 55% | -0.11 | 0.07 | 0.10 | 0.07 | 0.03 | -0.02 | 0.13 | 0.07 |
| 65% | -0.13 | 0.19 | 0.25 | 0.18 | 0.11 | 0.09 | 0.21 | -0.07 |
| 75% | -0.28 | 0.36 | 0.39 | 0.25 | 0.12 | 0.21 | 0.36 | -0.21 |
| 85% | -0.24 | 0.35 | 0.47 | 0.16 | 0.18 | 0.25 | 0.25 | -0.42 |
| 95% | 0.11 | 0.21 | 0.05 | -0.13 | 0.35 | 0.34 | 0.19 | -0.75 |
| 100% | 0.53 | -0.33 | -0.64 | -0.63 | 0.35 | 0.26 | -0.12 | -0.77 |
Clustering was k-means on the resampled z-curves, downsampled to 25 points per axis, with k chosen by mean silhouette from 2 to 6. The peak-end test used Spearman correlation against each film's rank. I do not have IMDb's ratings in the dataset, but the Top 250 is ordered by weighted rating, so rank is a monotone transform of rating and a rank correlation on one is identical to a rank correlation on the other.
What this cannot tell you
Dialogue measures what the characters say, not what the audience feels. The scariest stretches of a horror film have no dialogue at all, and a score cue can flip a scene the words did not. Silence is recorded in this data as an empty window and interpolated across, which is the right call for shape and the wrong call if you care about the silence itself.
The rank correlations are small and the range is restricted. Every film here is already in the top 250 of a database of millions. Finding a +0.22 inside that group is interesting. It is not a recipe.
The scores are a model's judgments. I have not hand-coded a validation set against them, and the sniff test in Chart 10 is not a validation set. It is a smell. The axes are also my choice, anchored to theories I find useful, and a different codebook could find shapes that this one cannot see. The clustering result in Chart 6 should be read with that in mind: no types on these axes.
And five-minute windows hide anything shorter than five minutes. A single devastating line inside an otherwise calm window gets averaged into the calm.
What I would do next is add a second rater and see where they disagree, and then go looking for the films that break the average curves on purpose, because the interesting ones always are. The practical version of all this, for people who make videos rather than films, is in the follow-up.
- Dr. J