Skip to main content
Back to all posts
Oct. 202612 min read

Alberta stopped changing its clocks in June. 19 of 19 frontier models still put Calgary on standard time in November.

On 18 June 2026 Alberta stopped changing its clocks. British Columbia had done it in March, the Northwest Territories followed in August, Morocco went back to plain UTC in September. All of it is in the database that runs the clock on your phone, and none of it is in a model trained before mid-2026. So I asked 19 models on Kaggle Benchmarks the question someone in Calgary is actually asking: it is 9 a.m. here on 15 November, what time is that in Toronto? The answer is 10:00 and all 19 said 11:00. Then I scored the same answer sheet against every tzdata release since 2022, which dates each model's world clock to the month it stopped.

Kaggle BenchmarksDEV ChallengeIANA tzdataKnowledge cutoffLLM evaluationTool useGPT-6 AstraGeminiClaudeHackathon
I wrote this post for the DEV x Kaggle Benchmarking Challenge (23 September to 11 October 2026), and built World Clock for the same submission. The benchmark is public on Kaggle as world-clock, made of two tasks, from memory and with a tzdata tool. Code, answer key, every result and both charts are at github.com/JonathanSolvesProblems/world-clock, open with no key and no account. #kagglechallenge
The claim in one sentence: across 122 questions graded against IANA tzdata 2026d, a public database I did not write and cannot influence, 13 of 19 frontier models got none of the 20 answers that changed in 2026, and all 19 put Calgary on the wrong offset in November. The grader also dates them: scoring the same answer sheet against all twenty tzdata releases since 2022 puts GPT-6 Astra's clock at the April 2026 release, against a vendor-stated cutoff of 30 April 2026, from nothing but clock questions. What is not claimed: that the dating is exact for every model. A model agreeing with its peak release on 46 of 46 changed answers has a sharp clock; one at 34 of 46 has a blurry one and its date is a best fit, so the table publishes both numbers.
A struck-through 11:00 beside a correct 10:00, labelled 19 of 19 models asked from memory on Kaggle Benchmarks against the IANA tz database 2026d, with a note that Alberta stopped changing its clocks on June 18.
The whole benchmark in one picture. It is 9 a.m. in Calgary on 15 November 2026. Toronto is 10:00. Every model asked said 11:00, and most cited the same rule: daylight saving ends on the first Sunday in November. That rule was true in Alberta for 55 years.

An answer key nobody has to write

Every model's knowledge stops at its training cutoff, and most of the time you cannot see the edge. Time zones are different. Governments change them by statute, on a date, and the IANA time zone database records every change within days. That gives a benchmark two things most benchmarks never get: an answer key nobody has to write, and a calendar to hold the model's knowledge against. The key here is tzdata 2026d, the September release of the same database that runs the clock on your phone, and I did not type a single expected answer. A script asks Python's zoneinfo what the clock did in a given place on a given date, and that is the truth. The maintainers ship a few releases a year, so the key rewrites itself and the benchmark does not go stale. 125 questions across six families: textbook controls like New York and Tokyo, awkward offsets like Kathmandu at +05:45 and Lord Howe Island's half-hour daylight saving, southern-hemisphere seasons, changes legislated between 2022 and 2025 in Iran, Jordan, Mexico, Kazakhstan and Chile, and the 2026 wave. Answers come back as structured output and grading is a string comparison after normalising offsets, so a wrong hour is a wrong hour and there is no judge model anywhere in the loop.

Two bar charts side by side. On the left, share of textbook questions right from memory, with every model near the top of the scale. On the right, changed 2026 answers right out of 20, where almost every bar is zero.
Time zones as a topic are learned. Fifteen of the 19 answered all 26 control questions correctly and nobody scored below 23. Time zones as of a date are a different thing entirely, which is the gap the right-hand chart measures.

Nobody knows about Alberta

Twenty of the 25 questions about the 2026 wave have an answer that changed this year. Thirteen of the 19 models got none of those 20 right, including Claude Opus 5, GPT-5.5, GPT-5.6 Terra and three of the five Geminis. Across all 19 models there were 18 correct answers on those 20 questions, and GPT-6 Astra produced 5 of them, four about British Columbia with the right reason attached. I read every one of the other 13 correct answers, and not one of those notes says that anything changed in 2026: they are right by accident. Qwen said the clocks do not change on 1 November because the fall-back is on 2 November, which is not a Sunday. Gemini 2.5 Pro said Inuvik's clocks do not change because the Northwest Territories observes Mountain Daylight Time year-round, then put Inuvik on the other offset in the next question. The Calgary offset question is the one I would put in front of a judge, because it cannot be right by accident: Calgary at noon on 15 November 2026 is -06:00, and all 19 models said -07:00.

A grid of 19 model names, each with the count of changed 2026 answers it got right out of 20. Most read zero, with a line underneath reading 13 of 19 got none of them.
Thirteen zeroes. The highest score on the board is 5 of 20, from the newest model in the set, and the 2025 models are scattered through the middle rather than at the bottom, which is its own finding.

Dating the clocks

This is the part the whole benchmark was built for. Of the 122 graded questions, 46 have an answer that changed between one tzdata release and another. I keep every release since 2022a unpacked, twenty of them, score each model's answer sheet against all twenty, and the release it agrees with most is the month its world clock stopped. GPT-6 Astra is the cleanest result in the set. Its curve rises through every release since 2022, agrees with tzdata 2026b on 45 of the 46 changed answers, and falls off a cliff at the next one. tzdata 2026b was released on 22 April 2026 and is the first to carry British Columbia; the release after it came on 8 July. So the ladder says this clock stopped between those two dates, and OpenAI states the model's cutoff as 30 April 2026. Those agree, and the ladder got there from nothing but questions about what time it is. The same method puts GPT-5.5, GPT-5.6 Terra and Claude Opus 5 on one shared curve at the last release before British Columbia, a perfect 46 of 46 each.

A line chart of answers matching each tzdata release out of the 46 that changed, with GPT-6 Astra's curve highlighted, rising to a peak at tzdata 2026b released April 22 2026 and dropping sharply after it, annotated with OpenAI's stated cutoff of April 30 2026.
The ladder. Everything before the peak right, everything after it wrong, and the peak lands four days before the vendor's published cutoff. Nothing about this chart knows what a cutoff is; it only knows what time the model thinks it is in twenty places.

The other result in this chart is the one I did not expect. Gemini 2.5 Pro from June 2025, Gemini 3.1 Pro from February 2026, Gemini 3.7 Flash from August 2026 and Gemini 3.8 Flash from September 2026 all peak at the same release, tzdata 2025a from 15 January 2025, and the two Flash curves never differ by more than one answer anywhere on the ladder. Google's own model card for 3.8 Flash gives a cutoff of March 2026 for some domains and January 2025 for the rest. The clock is one of the rest, and it has not moved in fifteen months of model releases. Release date is not the clock, and that is worth saying plainly because release date is what most people reach for when they want to know how current a model is.

The same ladder chart with three Gemini curves highlighted, Gemini 2.5 Pro, 3.7 Flash and 3.8 Flash, all peaking together at tzdata 2025a from January 2025.
Three models shipped fifteen months apart, one clock. The June 2025 model and the September 2026 model peak at the same January 2025 release.

Checking is not believing

Then I asked a subset again with a tool on the table. One function, zone_clock, which reads tzdata 2026d and returns the offset in force. The prompt says the tool exists and that the model may call it or answer without it. Nothing tells the model its knowledge might be out of date. Almost everyone checks: eleven of the 16 models asked on every one of the 43 graded questions. Checking mostly works: fourteen of the 16 got at least 34 of 43 with the tool where nobody got more than 23 from memory, and Claude Haiku 4.5, whose clock dates to 2022, got all 43 and never once answered against the database. But checking is not believing. On the question this post opens with, with the database one call away, 8 of the 16 still said 11:00, and two of those wrote the correct offsets in the same sentence. GPT-5.6 Terra's own note reads that Calgary used -06:00 and Toronto used -05:00. Its answer was 11:00.

A grid of 16 models answering the question what is 09:00 in Calgary in Toronto, with the tzdata tool available, eight showing 10:00 in green and eight showing 11:00 in orange.
The same question, with the database one call away. Half the models still got it wrong, and across all 16 there were 47 answers where the model asked about the right place on the right date and then answered something else.

Claude Opus 5 argued with the tool in so many words. On Calgary's offset it wrote that the tool reports -06:00 with abbreviation CST, which does not match Alberta's actual rules, before coming round to the tool's answer. On Casablanca in December it did not come round: it answered +01:00 and noted that the lookup returned +00:00, which conflicts with this known rule. Gemini 3.5 Flash-Lite called the tool three times about Calgary's clocks on 1 November, was told they do not change, and answered that they fall back. My favourite failure is Coyhaique: Chile's Aysén region got its own zone in tzdata 2025b, so a model whose clock stopped before that does not know the zone exists, asks the database about Santiago instead, gets a true answer about the wrong place, and reports it. Nine of the 16 missed that question and only five asked about the right zone at all. A lookup tool fixes what the model knows to look up.

A pull quote from Claude Opus 5 after calling the tool about Calgary, reading that the tool reports minus 06:00 with abbreviation CST, which does not match Alberta's actual rules.
A model telling the time zone database it is wrong about Alberta. It came round on this one. On Casablanca it did not.

Manitoba, which the database caught up with

Manitoba announced on 17 September that it will not fall back on 1 November. When I built the answer key no release carried it, so its three questions were asked, recorded, and left out of the score. Then on 30 September the maintainers shipped tzdata 2026e with Manitoba in it. So I graded those three answers against 2026e separately. Winnipeg at noon on 15 November is -05:00, and 9 a.m. in Winnipeg is 9 a.m. in Toronto. All 19 models said -06:00 and all 19 said 10:00. Asked whether Winnipeg's clocks change on 1 November, 18 said they fall back, and the one that said they do not only got there by believing the fall-back is on 2 November. The benchmark ran before the database knew, and the database caught up nine days later. The models will take a training run.

The World Clock benchmark page on Kaggle, showing its description, a leaderboard across two tasks and 19 models, and a score against total cost chart.
The benchmark is public on Kaggle as two tasks, from memory and with the tzdata tool. Anyone can run a new model against the same 125 questions, and the answer key updates itself every time the maintainers ship a release.

What is not claimed

The dating is not exact for every model. A model agreeing with its peak release on 46 of 46 changed answers has a sharp clock and the date is a fact about its answers; one at 34 of 46 has a blurry clock and the date is a best fit, which is why the table publishes both numbers for every model. I also fixed the harness several times along the way, so 15 models ran the from-memory half at least twice, 42 runs in all: the median model's score moved by 2 questions between runs, the largest swing was 10, and the dated release held in 10 of the 15. Every run of every model still put Calgary on -07:00. Manitoba's three questions are excluded from the headline score by construction and reported separately. GLM-5 and DeepSeek-R1 each lost a call to their backends and are scored on 121 rather than 122, and Gemma 4 left six answers empty, which count as wrong. The tool half is a 46-question subset, not the full set, because Kaggle allows $10 of model quota a day and a tool loop costs several times a plain answer; my first full lineup drained it and every queued run failed on a 403. GPT-6 Astra is absent from that half because its API refuses function tools unless reasoning is off and refuses to turn reasoning off, and three models the proxy lists cannot be served at all, which I report rather than quietly drop. Finally, this measures one narrow kind of knowledge. It says nothing about whether a model is good, only about what month its world clock stopped.

Related project

World Clock: a Kaggle benchmark that asks 19 frontier models what time it is, grades every answer against the IANA tz database, and dates each model's clock to the month it stopped

View the project