No es la primera vez que hablamos de machine learning en ArqueoTimes (Luengo, 2023), pero quizás sí la primera vez que nos centramos en un uso tan específico como es su aplicación en la epigrafía. Como ya dimos a entender en ese anterior artículo, y citando a Humby y Palmer (2006):
«Data is the new oil. It’s valuable, but if unrefined it cannot really be used. It has to be changed into gas, plastic, chemicals, etc., to create a valuable entity that drives profitable activity; so must data be broken down, analyzed for it to have value» (Los datos son el nuevo petróleo. Son valiosos, pero si no se refinan, en realidad no se pueden usar. Tienen que convertirse en gasolina, plásticos, productos químicos, etc., para crear una entidad valiosa que impulse una actividad rentable; así también los datos deben descomponerse y analizarse para que tengan valor).
Figura 1. Esta inscripción (Inscriptiones Graecae, volumen 1, edición 3, documento 4, cara B (IG I3 4B)) registra un decreto relacionado con la Acrópolis de Atenas y data del 485/4 a.C. Fuente: Marsyas, Epigraphic Museum, WikiMedia Licencia: CC BY 2.5.
En el contexto de la epigrafía esta cita de Humby y Palmer cobra enorme relevancia. La epigrafía se encuentra constantemente con desafíos tales como la fragmentación o deterioro de las inscripciones y es aquí donde el algoritmo Ithaca elaborado en los laboratorios de Google DeepMind por un equipo liderado por Yannis Assael entra en juego, proporcionando una herramienta avanzada para restaurar y analizar inscripciones antiguas, incluso cuando estas están dañadas e incluso incompletas. Sin embargo es importante destacar que Ithaca no es el único ni el primer ejemplo de uso de machine learning en epigrafía o sobre textos antiguos: Kang et al. (2021) emplearon modelos de lenguaje neuronal y técnicas de traducción automática para restaurar y analizar los registros de la dinastía Joseon. Otros como Bamman y Burns (2021) desarrollaron Latin BERT, un modelo de lenguaje contextual específicamente entrenado para el latín, capaz de predecir texto faltante demostrando la amplia aplicabilidad de estas tecnologías en la preservación y estudio de textos históricos en diversos contextos culturales.
¿Pero qué es Ithaca? Descripción del algoritmo
Imaginemos que tenemos un puzzle en el que las piezas están dañadas, desgastadas o incluso faltan por completo. El trabajo de los epigrafistas en este contexto es como el de un restaurador de puzzles: tratan de encajar las piezas que quedan y reconstruir la imagen original basándose en pistas dadas por el contexto. El algoritmo Ithaca actúa como un complemento que, en lugar de intentar simplemente encajar piezas visibles, puede inferir la forma y el contenido de las piezas que faltan basándose en patrones y experiencias aprendidas de miles de puzzles similares (en este caso, inscripciones antiguas).
El equipo subraya que Ithaca no está destinada a reemplazar al científico, sino a servir como una poderosa herramienta de apoyo. De hecho, el estudio enfatiza que mientras un historiador por sí solo alcanza una precisión del 25% y el algoritmo Ithaca por separado llega a un 62%, la combinación de ambos, historiador y algoritmo, logra un impresionante 72% de precisión.
Pero no podemos quedarnos en esos datos sin profundizar algo más porque, ¿qué quiere decir un 72% de precisión en este contexto, cómo funciona este algoritmo realmente y por qué está revolucionando la epigrafía?
Datos y Modelo del Algoritmo Ithaca
Para comprender cómo Ithaca alcanza su nivel de precisión, es crucial entender tanto los datos utilizados en el entrenamiento como la arquitectura del algoritmo. El trabajo, presentado a la comunidad científica el 9 de marzo de 2022 en la revista Nature (Assael et al., 2022), introdujo a Ithaca como una red neuronal profunda (deep neural network) diseñada para ayudar en la transcripción de fragmentos de texto hasta ahora ilegibles y en la identificación de la ubicación original y la fecha de datación de inscripciones antiguas, particularmente de la antigua Grecia. Para entrenar el modelo, los investigadores partieron de un corpus epigráfico compuesto por 178.551 inscripciones del Packard Humanities Institute (PHI), de las cuales 78.608 fueron seleccionadas para el entrenamiento, convirtiéndose en el dataset de texto epigráfico más grande manejado por machine learning hasta la fecha. Este conjunto de inscripciones están escritas en griego antiguo y han sido halladas a lo largo de todo el Mediterráneo antiguo, con dataciones entre el siglo VII a.C. y el siglo V d.C.
Figura 2. La arquitectura de Ítaca procesando la frase ‘δήμο το αθηναίων’ (‘el pueblo de Atenas’). Fuente. Licencia: Creative Commons Attribution 4.0.
El diseño del algoritmo está basado en una arquitectura de transformadores o transformers (un transformador es una arquitectura de aprendizaje profundo desarrollada por investigadores de Google y basada en el mecanismo de atención según Vaswani et al., 2017). Este sistema se estructura en dos partes principales: un «body» de transformadores y tres «heads» específicas para cada tarea: restauración del texto, atribución geográfica, y atribución cronológica. Este enfoque permite al algoritmo considerar el contexto de forma amplia y las conexiones entre diferentes partes del texto, incluso si algunas partes están fragmentadas o dañadas. Para poder entender mejor estos conceptos utilizaremos la siguiente analogía: piensa en el body principal de Ithaca como una persona en una fiesta que escucha múltiples conversaciones a la vez. Esta persona puede enfocarse en la conversación más relevante mientras todavía capta fragmentos de otras conversaciones cercanas. Del mismo modo, el transformer de Ithaca se enfoca en ciertas partes del texto (palabras y caracteres) que son más relevantes para entender el contexto completo, mientras aún «escucha» el resto del texto que podría ser relevante para futuras interpretaciones. Este enfoque permite que el algoritmo considere el contexto amplio y las conexiones entre diferentes partes del texto, incluso si algunas partes están fragmentadas o dañadas.
Como entradas epigráficas, Ithaca utiliza representaciones combinadas de caracteres y palabras, abordando así la pérdida de partes del texto. Además, introduce un símbolo especial «[unk]» cuando existen palabras dañadas o desconocidas. Las salidas del body se canalizan hacia las tres heads, consistiendo cada una en una red neuronal prealimentada poco profunda (shallow feedforward neural network —shallow FNN—) entrenada específicamente para su tarea. Por ejemplo, la head de restauración predice caracteres faltantes, la de atribución geográfica clasifica la inscripción entre 84 regiones, y la de atribución cronológica estima la fecha en intervalos de 10 años entre el 800 a.C. y el 800 d.C. El modelo también genera visualizaciones interpretativas, como mapa de prominencia (saliency maps) y listas clasificadas de predicciones, que facilitan la colaboración entre el modelo y los historiadores.
Con respecto a la idea del 72% de precisión, en el contexto del modelo Ithaca, esto significa que en el 72% de los casos, la predicción más alta del modelo es la correcta. Como ya indicamos líneas arriba, Ithaca no es el primer algoritmo en realizar estas tareas. Pythia, otro modelo de aprendizaje automático que fue desarrollado por un equipo igualmente liderado por Assael (2019), también ha abordado estas tareas anteriormente y es parte interna fundamental del nuevo algoritmo. Pero como era de esperar, Ithaca mejora el desempeño en todas las áreas en comparación con Pythia. Por ejemplo, para la restauración de texto (Restoration), Ithaca presenta una menor tasa de error de carácter (CER) con un 18.3% frente al 47.0% de Pythia. Igualmente, en la opción Top-1 Ithaca (de forma independiente) es muy superior, ostentando un 61.8% frente al 32.6% de Pythia.
Figura 3. Resultados generados por el algoritmo Ithaca. Fuente. Licencia: Creative Commons Attribution 4.0.
Seguramente con el tiempo podamos ir digitalizando y engrosando los corpus epigráficos aumentando el número de datos de entrenamiento lo que permita a la larga mejorar los modelos. Pero al respecto y para terminar, quizás sea adecuado que hagamos una reflexión fundamental. La precisión de estos modelos se mide con respecto a lo que hemos etiquetado previamente. Es decir, son los investigadores los que etiquetan el set de entrenamiento. Una vez el modelo está entrenado el algoritmo es capaz de continuar infiriendo dataciones por sí solo. Sin embargo, si los datos de entrenamiento fueran erróneos, todo el entrenamiento generaría un modelo erróneo, que sólo podría ofrecer respuestas equivocadas. En conclusión, y parafraseando las ideas de Einstein sobre la ciencia: «Una teoría puede caer con una nueva evidencia, pero un dato erróneo puede enraizarse profundamente en el tejido del conocimiento» (inspirado en «Mis ideas y visión del mundo», de Albert Einstein (2023)).
Bibliografía
Assael, Y., Sommerschield, T., Shillingford, B., Bordbar, M., Pavlopoulos, J., Chatzipanagiotou, M., Androutsopoulos, I., Prag, J., & de Freitas, N. (2022). Restoring and attributing ancient texts using deep neural networks. Nature, 603(7900), 280-283. https://doi.org/10.1038/s41586-022-04448-z
Assael, Y., Sommerschield, T., & Prag, J. (2019). Restoring ancient text using deep learning: A case study on Greek epigraphy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 6368-6375).Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1668
Bamman, D., & Burns, P. J. (2021). Latin BERT: A contextual language model for classical philology. Manuscript in preparation. https://doi.org/10.48550/arXiv.2009.10053
Kang, K., et al. (2021). Restoring and mining the records of the Joseon dynasty via neural language modeling and machine translation. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL) (pp. 4031–4042). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.320
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. En Proceedings of the 31st International Conference on Neural Information Processing Systems (pp. 6000–6010). Curran Associates Inc.
This translation was generated automatically using qwen2.5:14b-instruct and has not been reviewed by a human editor. It may contain errors, omissions or awkward phrasing — for the authoritative version, please read the original in Spanish. Found a mistake? Let us know at info@arqueotimes.es.
This is not the first time we have talked about machine learning in ArqueoTimes (Luengo, 2023), but perhaps it is the first time that we focus on such a specific use as its application in epigraphy. As we hinted at in that previous article, and citing Humby and Palmer (2006):
«Data is the new oil. It’s valuable, but if unrefined it cannot really be used. It has to be changed into gas, plastic, chemicals, etc., to create a valuable entity that drives profitable activity; so must data be broken down, analyzed for it to have value» (Data is the new oil. It is valuable, but if not refined, it cannot actually be used. It needs to be converted into gasoline, plastics, chemicals, etc., to create a valuable entity that drives profitable activity; thus, data must be decomposed and analyzed to have value).
Figure 1. This inscription (Inscriptiones Graecae, volume 1, edition 3, document 4, side B (IG I3 4B)) records a decree related to the Acropolis of Athens and dates from 485/4 BC. Source: Marsyas, Epigraphic Museum, WikiMedia License: CC BY 2.5.
In the context of epigraphy, this quote from Humby and Palmer takes on enormous relevance. Epigraphy constantly faces challenges such as the fragmentation or deterioration of inscriptions, and it is here where the Ithaca algorithm developed in Google DeepMind's labs by a team led by Yannis Assael comes into play, providing an advanced tool for restoring and analyzing ancient inscriptions, even when they are damaged or incomplete. However, it is important to note that Ithaca is not the only nor the first example of using machine learning in epigraphy or on ancient texts: Kang et al. (2021) employed neural language models and machine translation techniques to restore and analyze records from the Joseon dynasty. Others like Bamman and Burns (2021) developed Latin BERT, a contextual language model specifically trained for Latin, capable of predicting missing text demonstrating the wide applicability of these technologies in preserving and studying historical texts across various cultural contexts.
But what is Ithaca? Description of the algorithm
Imagine that we have a puzzle where pieces are damaged, worn out, or even completely missing. The work of epigraphers in this context is like that of a puzzle restorer: they try to fit together the remaining pieces and reconstruct the original image based on clues provided by the context. The Ithaca algorithm acts as an add-on that, instead of simply trying to fit visible pieces, can infer the shape and content of missing pieces based on patterns and experiences learned from thousands of similar puzzles (in this case, ancient inscriptions).
The team emphasizes that Ithaca is not meant to replace the scientist but to serve as a powerful support tool. In fact, the study highlights that while an historian alone achieves 25% accuracy and the Ithaca algorithm separately reaches 62%, the combination of both, historian and algorithm, achieves a remarkable 72% accuracy.
But we cannot stop at these data without delving deeper because, what does a 72% accuracy mean in this context, how does this algorithm really work, and why is it revolutionizing epigraphy?
Data and Model of the Ithaca Algorithm
To understand how Ithaca achieves its level of accuracy, it is crucial to understand both the data used in training as well as the architecture of the algorithm. The work, presented to the scientific community on March 9, 2022, in Nature (Assael et al., 2022), introduced Ithaca as a deep neural network designed to help transcribe fragments of text previously illegible and identify the original location and dating of ancient inscriptions, particularly from ancient Greece. To train the model, researchers started with an epigraphic corpus composed of 178,551 inscriptions from the Packard Humanities Institute (PHI), of which 78,608 were selected for training, becoming the largest dataset of epigraphic text managed by machine learning to date. This set of inscriptions is written in ancient Greek and have been found throughout the ancient Mediterranean, with dates ranging from the 7th century BC to the 5th century AD.
Figure 2. The architecture of Ithaca processing the phrase ‘δήμο το αθηναίων’ (‘the people of Athens’). Source. License: Creative Commons Attribution 4.0.
The design of the algorithm is based on a transformer architecture (a transformer is a deep learning architecture developed by Google researchers and based on the attention mechanism, according to Vaswani et al., 2017). This system is structured into two main parts: a «body» of transformers and three specific «heads» for each task: text restoration, geographic attribution, and chronological attribution. This approach allows the algorithm to consider context broadly and connections between different parts of the text, even if some parts are fragmented or damaged. To better understand these concepts, we will use the following analogy: think of Ithaca's main body as a person at a party listening to multiple conversations simultaneously. This person can focus on the most relevant conversation while still picking up fragments of nearby conversations. Similarly, Ithaca's transformer focuses on certain parts of the text (words and characters) that are more relevant for understanding the complete context, while still «listening» to the rest of the text that might be relevant for future interpretations. This approach allows the algorithm to consider broad context and connections between different parts of the text, even if some parts are fragmented or damaged.
As epigraphic inputs, Ithaca uses combined character and word representations, addressing thus the loss of parts of the text. Additionally, it introduces a special symbol «[unk]» when there are damaged or unknown words. The outputs from the body are channeled to the three heads, each consisting of a shallow feedforward neural network (shallow FNN) trained specifically for its task. For example, the restoration head predicts missing characters, the geographic attribution head classifies the inscription among 84 regions, and the chronological attribution head estimates the date in intervals of ten years between 800 BC and AD 800. The model also generates interpretative visualizations such as saliency maps and ranked prediction lists that facilitate collaboration between the model and historians.
Regarding the idea of 72% accuracy, in the context of the Ithaca model, this means that in 72% of cases, the highest prediction by the model is correct. As we indicated above, Ithaca is not the first algorithm to perform these tasks. Pythia, another machine learning model developed by a team also led by Assael (2019), has addressed these tasks previously and forms an integral part of the new algorithm. But as expected, Ithaca improves performance in all areas compared to Pythia. For example, for text restoration (Restoration), Ithaca presents a lower character error rate (CER) with 18.3% versus 47.0% of Pythia. Similarly, in the Top-1 option, Ithaca (independently) is far superior, boasting 61.8% against 32.6% for Pythia.
Figure 3. Results generated by the Ithaca algorithm. Source. License: Creative Commons Attribution 4.0.
Surely over time we can digitize and enrich epigraphic corpora, increasing the number of training data which will allow us to improve models in the long run. But regarding this and to conclude, perhaps it is appropriate for us to make a fundamental reflection. The accuracy of these models is measured relative to what has been previously labeled. That is, researchers label the training set. Once the model is trained, the algorithm can continue inferring dates on its own. However, if the training data were erroneous, all training would generate an erroneous model that could only offer incorrect answers. In conclusion, and paraphrasing Einstein's ideas about science: «A theory may fall with new evidence, but a wrong datum may become deeply rooted in the fabric of knowledge» (inspired by «My Ideas and Vision of the World», Albert Einstein (2023)).
Bibliography
Assael, Y., Sommerschield, T., Shillingford, B., Bordbar, M., Pavlopoulos, J., Chatzipanagiotou, M., Androutsopoulos, I., Prag, J., & de Freitas, N. (2022). Restoring and attributing ancient texts using deep neural networks. Nature, 603(7900), 280-283. https://doi.org/10.1038/s41586-022-04448-z
Assael, Y., Sommerschield, T., & Prag, J. (2019). Restoring ancient text using deep learning: A case study on Greek epigraphy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 6368-6375). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1668
Bamman, D., & Burns, P. J. (2021). Latin BERT: A contextual language model for classical philology. Manuscript in preparation. https://doi.org/10.48550/arXiv.2009.10053
Kang, K., et al. (2021). Restoring and mining the records of the Joseon dynasty via neural language modeling and machine translation. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL) (pp. 4031–4042). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.320
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (pp. 6000–6010). Curran Associates Inc.
Luengo Gutiérrez, F. J. (2024, 16 de septiembre). El algoritmo Ithaca. DeepMind al servicio de la Arqueología. ArqueoTimes. https://arqueotimes.com/articulos/el-algoritmo-ithaca-deepmind-al-servicio-de-la-arqueologia/
Comentarios