Pausanias Analysis

Pipeline Progress

Task Progress

AreaTaskScriptCadenceDoneTotalProgressStatusEst. completionDetails
Core labels Passage mythic/skeptic labels mythic_sceptic_analyser.py 50/day 3,170 3,170
100.0%
Complete Done Passage-level mythic-era and scepticism flags.
Names and places Proper-noun extraction extract_proper_nouns.py 50/day 3,170 3,170
100.0%
Complete Done Sections processed for proper nouns.
Names and places Wikidata linking link_wikidata.py 100/day 6,733 6,733
100.0%
Complete Done Unique name/type pairs linked or reviewed.
Text preparation Translation translate_pausanias.py 50/day 3,170 3,170
100.0%
Complete Done Passage-level English translations.
Text preparation Sentence splitting split_sentences.py 200/day 3,170 3,170
100.0%
Complete Done 11,302 aligned Greek/English sentences available.
Text preparation Passage summaries summarise_passages.py 50/day 3,170 3,170
100.0%
Complete Done Short summaries used on reader pages.
Statistical models Passage predictor models find_predictors.py refresh 4 4
100.0%
Complete Done Full and simplified passage-level TF-IDF models.
Statistical models Sentence predictor models find_sentence_predictors.py refresh 4 4
100.0%
Complete Done Sentence-level TF-IDF model artifacts.
Names and places Proper-noun network analyse_noun_network.py refresh 6,733 6,733
100.0%
Complete Done Co-occurrence network and centrality rows.
Sentence labels Greta-inspired sentence tags sentence_tagging_daily.sh Batch API, ~2,205/day 11,302 11,302
100.0%
Complete Done completed: 1
Sentence labels Original sentence tags sentence_tagging_daily.sh Batch API, ~917/day 11,302 11,302
100.0%
Complete Done completed: 13
Sentence labels Legacy sentence classifier sentence_tagging_daily.sh 5/day comparison 5,751 11,302
50.9%
In progress 2029-08-04 completed: 61; completed with failures: 1
Grammar LLM grammar analysis sentence_llm_grammar_daily.sh 1M tokens/day, ~424/day 11,301 11,302
100.0%
In progress 2026-07-21 completed: 2; completed budget exhausted: 1; completed with failures: 4; completed with failures budget exhausted: 23; failed failure limit: 2; running: 3
Sentence labels Discourse mode sentence tags sentence_tagging_daily.sh Batch API, 100k tokens/day 2,400 11,301
21.2%
In progress 2026-10-18 batch submitted: 1; completed: 24
Grammar UDPipe grammar analysis sentence_udpipe.py parser pass 11,302 11,302
100.0%
Complete Done completed: 2
Grammar Trankit grammar analysis sentence_trankit.py parser pass 11,302 11,302
100.0%
Complete Done completed: 3
Lemmas Sentence lemmatization sentence_lemmatizer.py auxiliary LLM pass 100 11,302
0.9%
In progress Unknown completed: 1
Lemmas Word-form lemmatization word_lemmatizer.py Batch API/ad hoc 28,580 n/a n/a Tracking n/a batch in progress: 1; completed: 3
Names and gender People/name analysis section_people_daily.sh Batch API, ~40 sections/day 960 3,170
30.3%
In progress 2026-09-14 batch submitted: 1; completed: 24; 6,937 mention rows

Token Usage

SourceInput tokensOutput tokensTotal
Passage analysis 1,489,852 57,936 1,547,788
Noun extraction 2,592,448 1,041,140 3,633,588
Translation 1,233,082 379,009 1,612,091
Summarisation 758,681 1,303,428 2,062,109
Phrase translation 134,137 802,284 936,421
Sentence tagging 16,543,648 1,769,993 18,313,641
LLM grammar 8,800,418 17,802,800 26,603,218
Sentence lemmatization 48,105 10,244 58,349
Word lemmatization 310,141 801,862 1,112,003
Section people extraction 1,043,207 473,725 1,516,932
Total 32,953,719 24,442,421 57,396,140