·11 min read

I fed 195 of the best films ever made to a model and charted where the feeling goes

Ten things the dialogue of the IMDb Top 250 says about how great films handle jokes, tension, endings and grief, each with a chart. Then the method, the caveats, and what did not work.

I had a folder in Google Drive with the English subtitles for the IMDb Top 250 and no plan for it. The idea I kept coming back to is from granular synthesis. You take a sound, cut it into tiny overlapping grains, and the one continuous thing becomes a pile of small things you can measure and rearrange. Do that to a film and you get a few dozen five-minute windows of dialogue, each small enough to ask a simple question about. How tense is this? How funny? Is anybody grieving?

So I did it. 195 films, 10,129 windows, 3.8 million words of dialogue, eight emotion scores per window plus a dominant emotion, all placed on a shared clock that runs from 0% to 100% of the film so a 90-minute comedy and a three-hour epic can be laid on top of each other. The ten things I found come first. The method, the caveats, and the stuff that did not work are after that, for the people who want to argue.

Ten things the charts say

1. Great films front-load the jokes

Humor peaks about a quarter of the way in and sinks below the film's own average by the end. This is dialogue humor only, so a sight gag or a Buster Keaton fall is invisible here. The words still tell one consistent story: you get the laughs early and you pay for them later.

Chart 1

Humor peaks early and never comes back

Mean humor score across all films at each point of narrative time. Each film is z-scored first, so this is shape, not level. The band is one standard deviation.

-1010%25%50%75%100%z-score, mean of all films (band: ±1 SD)share of runtime →peak at 23%
Table view
runtimehumor
5%0.16
14%0.37
25%0.44
35%0.21
45%0.22
55%0.07
65%-0.07
75%-0.21
85%-0.42
95%-0.75
100%-0.77
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

2. Tension climbs the whole film and then falls off a cliff

On average, tension rises for 88% of the runtime and then drops harder than anything else in the data. That drop is the resolution. Arousal does the same thing a beat behind. It is the most consistent shape in the whole set.

Chart 2

Tension rises for nearly the entire film, then collapses

Mean tension and arousal across all films, z-scored per film. The last few percent of runtime is where the corpus moves together.

-1010%25%50%75%100%z-score, mean of all filmsshare of runtime →tensionarousaltension peaks at 88%
Table view
runtimetensionarousal
5%-0.52-0.53
14%-0.26-0.19
25%-0.19-0.16
35%-0.08-0.10
45%-0.02-0.05
55%0.100.07
65%0.250.19
75%0.390.36
85%0.470.35
95%0.050.21
100%-0.64-0.33
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

3. The ending is the one place these films agree

For most of the runtime the films disagree about where their beats go, which is why the average curves are so flat in the middle. Then the last tenth arrives and they line up. Warmth peaks there in 31% of films and valence in 30%. Humor peaks there in 2%.

Chart 3

Where the peak lands: the share of films that save it for the last 10%

For each axis, the percentage of films whose curve reaches its maximum in the final tenth of runtime.

0%10%20%30%40%warmth31%valence30%grief23%moral weight22%tension22%conflict17%arousal16%humor2%
Table view
films peaking in last 10%
warmth31%
valence30%
grief23%
moral weight22%
tension22%
conflict17%
arousal16%
humor2%
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

4. Warmth, grief and moral weight all get heavier as the film goes

On the raw 0 to 4 scale, the average film gains about half a point of warmth, grief and moral weight between its first tenth and its last tenth. Humor loses more than half a point. Valence and conflict barely move on average, which hides a lot of individual movement. That is the next chart.

Chart 4

Opening level versus closing level, by axis

Average score in the first 10% of runtime against the last 10%, on the 0 to 4 rubric scale. Sorted by how much each axis rises.

01234first 10% of runtimelast 10% of runtimewarmth+0.51moral weight+0.47grief+0.46arousal+0.28tension+0.27valence+0.01conflict+0.00humor-0.64
Table view
axisfirst 10%last 10%change
warmth1.401.91+0.51
moral weight1.902.37+0.47
grief1.051.51+0.46
arousal2.522.80+0.28
tension2.212.47+0.27
valence1.481.50+0.01
conflict1.831.83+0.00
humor1.801.16-0.64
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

5. Fewer than half of them end happier than they start

93 of the 195 end with higher valence than they opened with. 102 end lower. The happy ending is not the default in this canon, and a lot of the films you would call uplifting get there through an ending that is quieter and sadder than their opening.

Chart 5

Ending valence minus opening valence, one bar per half point

Each film's last-10% average minus its first-10% average. Bars to the right of zero are films that end up.

49365343301523-2.5-2.0-1.5-1.0-0.50+0.5+1.0+1.5+2.0+2.5ending valence minus opening valence, 0 to 4 scale102 films end lower93 films end higher
Table view
change in valencefilms
-2.5 to -2.00
-2.0 to -1.54
-1.5 to -1.09
-1.0 to -0.536
-0.5 to 053
0 to +0.543
+0.5 to +1.030
+1.0 to +1.515
+1.5 to +2.02
+2.0 to +2.53
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

6. There is no secret set of story shapes

I clustered the curves looking for the six or seven shapes everyone hopes to find. The best split on all eight axes has a silhouette score of 0.04. On valence alone it is 0.08. A score of 1 would mean clean separation, and 0 means the clusters are no tighter than a random cut. The honest reading is that these films sit on a continuum. It could be that my axes are the wrong ones, and I say more about that below, but on this data there are tendencies, not types.

Chart 6

How separable the clusters are, for every number of clusters I tried

Mean silhouette from k-means on the z-scored curves. Nothing gets near the range where you would call the groups real.

0.00.10.20.30.40.5mean silhouette (1 = perfectly separated, 0 = no structure)all eight axesvalence onlyk = 20.040.08k = 30.040.08k = 40.030.07k = 50.030.07k = 60.030.06
Table view
kall eight axesvalence only
k = 20.0360.082
k = 30.0360.078
k = 40.0330.069
k = 50.030.066
k = 60.0310.063
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

7. The one split that does show up is "ends up" versus "ends down"

When you force the valence curves into two groups, the only thing the groups disagree about is the last few minutes. 90 films rise at the close. 105 fall. Everything before that looks the same.

Chart 7

The two valence clusters, and the only place they differ

Mean z-scored valence curve of each cluster. Identical for 90% of the runtime, then they part.

-1010%25%50%75%100%z-scoreshare of runtime →ends up (90)ends down (105)
Table view
runtimeends up (90)ends down (105)
5%0.160.31
14%-0.240.19
25%0.090.12
35%-0.230.30
45%-0.100.11
55%-0.18-0.05
65%-0.16-0.10
75%-0.22-0.39
85%-0.11-0.36
95%0.89-0.61
100%1.410.04
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

8. The peak-end rule does not predict which films rank highest

Kahneman's rule says we remember an experience by its most intense moment and its ending. I tested it. Correlate each film's valence peak, its ending, and the average of the two against its place on the list. All three land near zero. The one valence number that moves with rank is the opening, and it moves the wrong way for a feel-good theory: films that open darker rank higher.

Chart 8

Valence measures against IMDb rank

Spearman correlation between each measure and rank, with rank flipped so a positive number means more of this goes with a higher place on the list. Values above 0.14 in either direction clear p below .05, and are marked.

-0.3-0.2-0.10+0.1+0.2+0.3mean-0.12peak-0.04end-0.03peak end-0.05start-0.16 *
Table view
rho vs. rank
mean-0.12
peak-0.04
end-0.03
peak end-0.05
start-0.16 *
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

9. Tension and moral stakes go with a higher rank. Jokes go the other way.

Average tension across the whole film correlates with rank at +0.22, moral weight at +0.17. Humor is negative and not significant. These are modest numbers inside an already elite list, so read them as which way the wind blows, not as a formula. The wind blows toward stakes.

Chart 9

A film's average level on each axis, against its rank

Same test as Chart 8, one bar per axis, using the film's mean score across its whole runtime.

-0.3-0.2-0.10+0.1+0.2+0.3tension+0.22 *moral weight+0.17 *arousal+0.09conflict+0.07grief+0.05warmth+0.05humor-0.10valence-0.12
Table view
rho vs. rank
tension+0.22 *
moral weight+0.17 *
arousal+0.09
conflict+0.07
grief+0.05
warmth+0.05
humor-0.10
valence-0.12
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

10. The rater passed the sniff test

Before you trust any of the above you want to know the model reads dialogue the way a person would. The most positive film on the list is Singin' in the Rain. The funniest is Monty Python and the Holy Grail. The most tense is Dr. Strangelove. Downfall, Full Metal Jacket and Oldboy sit at the bottom of valence, and the loudest film in the set is The Wolf of Wall Street, which anyone who has seen it will accept without a chart.

Chart 10

The top three films on three axes

Highest average score across the whole film. If these looked wrong, nothing above would be worth reading.

most positiveSingin' in the Rain (1952)2.89WALL·E (2008)2.51My Neighbor Totoro (1988)2.48funniestMonty Python and the Ho... (1975)3.72Some Like It Hot (1959)3.16Finding Nemo (2003)3.09most tenseDr. Strangelove or- How... (1964)3.61Avengers- Infinity War (2018)3.49The Dark Knight (2008)3.47
Table view
filmmean level
most positive #1Singin' in the Rain (1952)2.89
most positive #2WALL·E (2008)2.51
most positive #3My Neighbor Totoro (1988)2.48
funniest #1Monty Python and the Holy Grail (1975)3.72
funniest #2Some Like It Hot (1959)3.16
funniest #3Finding Nemo (2003)3.09
most tense #1Dr. Strangelove or- How I Learned to Stop Worrying and Love the Bomb (1964)3.61
most tense #2Avengers- Infinity War (2018)3.49
most tense #3The Dark Knight (2008)3.47
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

How I did it

The subtitles came from a folder of 200 SRT files, one per film, ranked at the time they were collected. Five came out before any analysis: four silent films that only have intertitles (Modern Times, City Lights, Metropolis, The Kid) and Kill Bill: The Whole Bloody Affair, whose subtitle file has 230 cues for a 232-minute cut and is clearly incomplete. Kill Bill Vol. 1 stayed. That leaves 195.

Each file was parsed, stripped of formatting tags and the ads subtitle sites inject, and cut into windows.

Chart 11

Five-minute windows, every two and a half minutes

The grain size. Two minutes left too many windows with almost no dialogue during action and montage. Ten minutes gave only a dozen points per film, too few to see a shape.

film runtimecredits5 min...every 2.5 min, so each window overlaps its neighbors by halfeach window's dialogue goes to the rater and comes back as 8 scores (0 to 4) plus a dominant emotion
Table view
settingvalue
window length5 min
hop2.5 min
overlap50%
windows per filmabout 50
The window with no dialogue at all is recorded as empty and skipped, not scored as calm. Silence in a war film is not calm.

Every window's dialogue went to Jev, TypeSafe AI's evaluation model, through the Vercel AI Gateway, where it is free. Jev is not a chat model. You give it a piece of text and a set of typed questions, and it returns a calibrated answer for each without generating any prose. For a score question you supply the rungs of a rubric in order, and it returns a probability for each rung plus an interpolated score. That is a better fit for this job than asking a chat model to emit JSON and hoping the numbers mean the same thing on window 9,000 as they did on window 1. The model saw only the dialogue. No title, no year. If it knew it was reading The Godfather it would rate from memory of the film instead of from the text, and the measurement would be gone.

The eight axes, each a five-rung rubric scored 0 to 4, and the theory each one is anchored to:

| Axis | Anchored to | |---|---| | valence | Russell's circumplex model of affect | | arousal | Russell, and Zillmann's excitation transfer | | tension | Brewer and Lichtenstein's structural affect; Huron's theory of expectation | | conflict | Zillmann's affective disposition theory | | warmth | kama muta, the "being moved" research from Fiske and Menninghaus | | grief | Bonanno's grief trajectories; Herman's stages of recovery | | moral weight | Haidt's moral foundations; Shay's moral injury | | humor | its own axis, because humor is not the opposite of sadness |

Plus a ninth question that picks one dominant emotion from Plutchik's eight, or neutral. The pipeline is Bun and TypeScript, eight calls in parallel, and it finished 10,129 windows in about twenty minutes with 1,690 retries and no failures. The gateway was having a day. I had Claude Code write most of it.

Then the normalization. Every film's windows were placed on its own runtime, resampled to 100 points from 0% to 100%, and z-scored per film per axis. The z-score is the choice that makes the curves comparable. It throws away how tense a film is on average and keeps only where in the film the tension goes. The raw 0 to 4 levels are kept separately for the questions that need them.

Chart 12

All eight average curves

The same view as Charts 1 and 2 for every axis. Flat middles, busy endings.

valence1-10%100%arousal1-10%100%tension1-10%100%conflict1-10%100%warmth1-10%100%grief1-10%100%moral weight1-10%100%humor1-10%100%
Table view
runtimevalencearousaltensionconflictwarmthgriefmoral weighthumor
5%0.25-0.53-0.52-0.29-0.43-0.32-0.500.16
14%0.04-0.19-0.260.02-0.22-0.17-0.170.37
25%0.13-0.16-0.19-0.05-0.11-0.15-0.230.44
35%0.04-0.10-0.080.01-0.14-0.11-0.070.21
45%0.03-0.05-0.020.010.08-0.06-0.030.22
55%-0.110.070.100.070.03-0.020.130.07
65%-0.130.190.250.180.110.090.21-0.07
75%-0.280.360.390.250.120.210.36-0.21
85%-0.240.350.470.160.180.250.25-0.42
95%0.110.210.05-0.130.350.340.19-0.75
100%0.53-0.33-0.64-0.630.350.26-0.12-0.77
Source: 195 IMDb Top 250 films, English subtitle dialogue in 5-minute windows, scored by TypeSafe AI's Jev. Curves are on a 0 to 100% narrative clock.

Clustering was k-means on the resampled z-curves, downsampled to 25 points per axis, with k chosen by mean silhouette from 2 to 6. The peak-end test used Spearman correlation against each film's rank. I do not have IMDb's ratings in the dataset, but the Top 250 is ordered by weighted rating, so rank is a monotone transform of rating and a rank correlation on one is identical to a rank correlation on the other.

What this cannot tell you

Dialogue measures what the characters say, not what the audience feels. The scariest stretches of a horror film have no dialogue at all, and a score cue can flip a scene the words did not. Silence is recorded in this data as an empty window and interpolated across, which is the right call for shape and the wrong call if you care about the silence itself.

The rank correlations are small and the range is restricted. Every film here is already in the top 250 of a database of millions. Finding a +0.22 inside that group is interesting. It is not a recipe.

The scores are a model's judgments. I have not hand-coded a validation set against them, and the sniff test in Chart 10 is not a validation set. It is a smell. The axes are also my choice, anchored to theories I find useful, and a different codebook could find shapes that this one cannot see. The clustering result in Chart 6 should be read with that in mind: no types on these axes.

And five-minute windows hide anything shorter than five minutes. A single devastating line inside an otherwise calm window gets averaged into the calm.

What I would do next is add a second rater and see where they disagree, and then go looking for the films that break the average curves on purpose, because the interesting ones always are. The practical version of all this, for people who make videos rather than films, is in the follow-up.

  • Dr. J