Práctica · Edición científica

Auditoría bibliográfica con LLM

Fichero .bib con errores deliberados, prompt de revisión para Claude y versión corregida. Copia cada bloque con un clic.

1

Fichero .bib con errores deliberados

Este fichero contiene 18 entradas BibTeX del ámbito de Translation Studies con 14 errores incrustados de los tipos que Zotero, Mendeley y EndNote no detectan automáticamente: duplicados semánticos, tipologías de entrada incorrectas, campos faltantes, variantes de nombre de autor, títulos con erratas, DOIs rotos e inconsistencias en nombres de revistas.

18
Entradas totales
14
Errores incrustados
3
Duplicados semánticos
14
Entradas reales tras limpieza
referencias_con_errores.bib
%% ═══════════════════════════════════════════════════════════════ %% FICHERO DE PRÁCTICA — Auditoría bibliográfica con LLM %% 18 entradas · 14 errores deliberados · Translation Studies %% Fuente: Agentes IA en Humanidades · UGR %% ═══════════════════════════════════════════════════════════════ %% ── 1. Entrada correcta (referencia) ────────────────────────── @article{baker2018routledge, author = {Baker, Mona}, title = {In Other Words: A Coursebook on Translation}, journal = {Routledge}, year = {2018}, edition = {3rd}, address = {London}, doi = {10.4324/9781315619187} } %% ── 2. ERROR: tipo incorrecto (debería ser @book) ──────────── @article{munday2016introducing, author = {Munday, Jeremy}, title = {Introducing Translation Studies: Theories and Applications}, publisher = {Routledge}, year = {2016}, edition = {4th}, address = {London} } %% ── 3. Entrada correcta ────────────────────────────────────── @article{castilho2017machine, author = {Castilho, Sheila and Moorkens, Joss and Gaspari, Federico and Calixto, Iacer and Tinsley, John and Way, Andy}, title = {Is Neural Machine Translation the New State of the Art?}, journal = {The Prague Bulletin of Mathematical Linguistics}, year = {2017}, volume = {108}, number = {1}, pages = {109--120}, doi = {10.1515/praam-2017-0013} } %% ── 4. ERROR: DOI incorrecto (praam → pragl) ───────────────── %% ERROR: año incorrecto (2017 → 2018) ─────────────────────── %% ── 5. ERROR: duplicado semántico de #3 (variante de nombre) ─ @article{castilho2018neural, author = {Castilho, S. and Moorkens, J. and Gaspari, F. and Calixto, I. and Tinsley, J. and Way, A.}, title = {Is Neural Machine Translation the New State of the Art?}, journal = {Prague Bull. Math. Linguistics}, year = {2018}, volume = {108}, pages = {109--120}, doi = {10.1515/pragl-2017-0013} } %% ── 6. Entrada correcta ────────────────────────────────────── @article{toral2017multifaceted, author = {Toral, Antonio and Way, Andy}, title = {What Level of Quality Can Neural Machine Translation Attain on Literary Text?}, journal = {Translation Spaces}, year = {2018}, volume = {7}, number = {1}, pages = {81--106}, doi = {10.1075/ts.00007.tor} } %% ── 7. ERROR: falta año ────────────────────────────────────── @incollection{nord_functionalist, author = {Nord, Christiane}, title = {Translating as a Purposeful Activity: Functionalist Approaches Explained}, booktitle = {Translation Theories Explored}, publisher = {St. Jerome}, address = {Manchester} } %% ── 8. ERROR: errata en título ("Transltion" → "Translation") %% ERROR: campo journal en @inproceedings (debería ser booktitle) @inproceedings{koponen2016machine, author = {Koponen, Maarit}, title = {Machine Transltion Post-Editing and Effort: An Overview}, journal = {Proceedings of the 19th Annual Conference of the EAMT}, year = {2016}, pages = {131--142}, address = {Riga} } %% ── 9. Entrada correcta ────────────────────────────────────── @article{laubli2018machine, author = {Läubli, Samuel and Sennrich, Rico and Volk, Martin}, title = {Has Machine Translation Achieved Human Parity? A Case for Document-Level Evaluation}, journal = {Proceedings of EMNLP}, year = {2018}, pages = {4791--4796}, doi = {10.18653/v1/D18-1512} } %% ── 10. ERROR: duplicado semántico de #9 (Läubli → Laubli) ── %% ERROR: título ligeramente distinto ────────────────────── @article{laubli2018parity, author = {Laubli, Samuel and Sennrich, Rico and Volk, Martin}, title = {Has Machine Translation Achieved Human Parity? A Case for Document Level Evaluation}, journal = {Proc. of EMNLP 2018}, year = {2018}, pages = {4791--4796} } %% ── 11. ERROR: tipo incorrecto (@article → @book) ─────────── %% ERROR: campo "journal" no corresponde a un libro ──────── @article{hurtado2001traduccion, author = {Hurtado Albir, Amparo}, title = {Traducción y Traductología: Introducción a la traductología}, journal = {Cátedra}, year = {2001}, address = {Madrid} } %% ── 12. Entrada correcta ───────────────────────────────────── @article{moorkens2018translators, author = {Moorkens, Joss}, title = {What to Expect from Neural Machine Translation: A Practical In-Class Translation Evaluation Exercise}, journal = {The Interpreter and Translator Trainer}, year = {2018}, volume = {12}, number = {4}, pages = {375--387}, doi = {10.1080/1750399X.2018.1501639} } %% ── 13. ERROR: duplicado semántico de Baker #1 ─────────────── %% ERROR: variante de nombre y edición distinta ──────────── @book{baker2011other, author = {Baker, M.}, title = {In Other Words: A Coursebook on Translation}, publisher = {Routledge}, year = {2011}, edition = {2nd}, address = {London and New York} } %% ── 14. ERROR: falta publisher en @book ────────────────────── @book{venuti2012translation, author = {Venuti, Lawrence}, title = {The Translation Studies Reader}, year = {2012}, edition = {3rd}, address = {London} } %% ── 15. Entrada correcta ───────────────────────────────────── @article{guerberof2022creativity, author = {Guerberof-Arenas, Ana and Toral, Antonio}, title = {Creativity in Neural Machine Translation: A Preliminary Analysis}, journal = {Translation Spaces}, year = {2022}, volume = {11}, number = {1}, pages = {1--25}, doi = {10.1075/ts.21025.gue} } %% ── 16. ERROR: DOI inventado / inexistente ─────────────────── @article{kenny2020machine, author = {Kenny, Dorothy}, title = {Machine Translation for Everyone: Empowering Users in the Age of Artificial Intelligence}, journal = {Translation Studies}, year = {2020}, volume = {1}, pages = {1--16}, doi = {10.1080/14781700.2019.9999999} } %% ── 17. ERROR: falta pages en @article ─────────────────────── %% ERROR: nombre de revista abreviado inconsistente ──────── @article{obrien2012towards, author = {O'Brien, Sharon}, title = {Towards a Dynamic Quality Evaluation Model for Translation}, journal = {JoSTrans}, year = {2012}, volume = {17} } %% ── 18. Entrada correcta ───────────────────────────────────── @article{daems2017translation, author = {Daems, Joke and Vandepitte, Sonia and Hartsuiker, Robert and Desmet, Lieve}, title = {Translation Methods and Experience: A Comparative Analysis of Human Translation and Post-Editing}, journal = {Perspectives}, year = {2017}, volume = {25}, number = {3}, pages = {452--473}, doi = {10.1080/0907676X.2016.1246377} }

🔍 Catálogo detallado de los 14 errores

DUP-1#5 duplica a #3 — Castilho et al. Mismo artículo con iniciales vs. nombres completos, revista abreviada vs. completa, año cambiado (2017→2018)
DUP-2#10 duplica a #9 — Läubli et al. Diéresis eliminada, título sin guion en "Document-Level", revista con formato diferente, DOI ausente
DUP-3#13 duplica a #1 — Baker. Edición anterior (2ª vs. 3ª) del mismo libro, autor abreviado "Baker, M." vs. "Baker, Mona"
TIPO-1#2 Munday — Tipo @article pero es un @book (tiene publisher, edition, no tiene journal/volume/pages)
TIPO-2#11 Hurtado Albir — Tipo @article pero es un @book; campo "journal" contiene el nombre de la editorial (Cátedra)
TIPO-3#8 Koponen — @inproceedings usa campo "journal" en vez de "booktitle"
CAMPO-1#7 Nord — Falta el campo year (obligatorio en toda entrada BibTeX)
CAMPO-2#14 Venuti — Falta publisher en @book
CAMPO-3#17 O'Brien — Falta pages en @article; nombre de revista abreviado (JoSTrans vs. The Journal of Specialised Translation)
ERRATA-1#8 Koponen — "Transltion" → "Translation" en el título
DOI-1#5 Castilho (dup) — DOI incorrecto: pragl → praam (10.1515/pragl-2017-0013)
DOI-2#16 Kenny — DOI inventado: 10.1080/14781700.2019.9999999 no existe
AÑO-1#5 Castilho (dup) — Año cambiado de 2017 a 2018
NOMBRE-1#1 Baker — La entrada #1 es @article pero debería ser @book (3ª ed. de un manual)
2

Prompt de auditoría para Claude

Copia este prompt en Claude.ai (preferiblemente Opus para máxima precisión) y pega a continuación el contenido del fichero .bib. El modelo generará un informe de incidencias y la versión corregida completa.

Flujo de trabajo: (1) Copia el prompt → (2) Pégalo en Claude.ai → (3) Copia el .bib con errores → (4) Pégalo justo después del prompt → (5) Envía. El modelo devolverá el informe + .bib limpio en una sola respuesta.
Prompt · Auditoría bibliográfica exhaustiva
Eres un bibliotecario experto en gestión de referencias BibTeX para investigación en Translation Studies. Realiza una auditoría exhaustiva del fichero .bib que te proporcionaré a continuación. FASES DE LA AUDITORÍA: 1. DUPLICADOS SEMÁNTICOS Identifica entradas que refieren a la MISMA obra aunque difieran en: - Formato de nombre del autor (iniciales vs. nombre completo) - Variantes tipográficas del título (guiones, mayúsculas, diacríticos) - Nombre de revista completo vs. abreviado - Año ligeramente distinto en una de las copias Para cada duplicado: señala las dos entradas, explica por qué son la misma obra, indica cuál conservar (la más completa/correcta) y cuál eliminar. 2. TIPOLOGÍA DE ENTRADA Verifica que el tipo BibTeX (@article, @book, @incollection, @inproceedings...) sea coherente con los campos presentes: - Un @article DEBE tener journal, volume y pages - Un @book DEBE tener publisher - Un @inproceedings DEBE tener booktitle (NO journal) - Un @incollection DEBE tener booktitle y publisher Señala entradas mal tipificadas y propón el tipo correcto. 3. CAMPOS OBLIGATORIOS FALTANTES Para cada tipo, verifica la presencia de campos mínimos: - Todas: author, title, year - @article: + journal, volume, pages - @book: + publisher - @incollection: + booktitle, publisher - @inproceedings: + booktitle Lista los campos ausentes por entrada. 4. ERRATAS EN TÍTULOS Y DATOS Busca errores tipográficos en títulos, nombres de autores y otros campos de texto. 5. DOIs - Señala entradas sin DOI (post-2010 deberían tenerlo) - Detecta DOIs con formato sospechoso o claramente inventados - Verifica la estructura básica (10.XXXX/...) 6. CONSISTENCIA DE NOMBRES DE REVISTA Detecta si una misma revista aparece con distintas variantes (abreviada y completa, con y sin "The", etc.) y propón la forma canónica. FORMATO DE SALIDA: A) INFORME DE AUDITORÍA Tabla Markdown con columnas: | # | Clave BibTeX | Tipo de error | Gravedad | Descripción | Corrección propuesta | Gravedad: CRÍTICO (duplicado, tipo incorrecto) · MAYOR (campo faltante, DOI falso) · MENOR (errata, abreviatura) B) ESTADÍSTICAS - Total de entradas analizadas - Entradas con errores / sin errores - Duplicados eliminados - Entradas en el .bib limpio final C) FICHERO .BIB CORREGIDO Genera la versión limpia completa del fichero .bib con: - Duplicados eliminados (conservar la entrada más completa) - Tipos de entrada corregidos - Campos faltantes marcados con placeholder si no puedes inferirlos: year = {FALTA}, pages = {VERIFICAR} - Erratas corregidas - DOIs sospechosos marcados como: doi = {VERIFICAR: 10.xxxx/...} - Nombres de revista en su forma completa y canónica - Comentarios %% explicando cada corrección aplicada FICHERO .BIB A AUDITAR:
3

Versión corregida (resultado esperado)

Esta es la versión limpia que el modelo debería producir (o una muy similar). Úsala como referencia para evaluar la calidad de la auditoría generada. Contiene 14 entradas únicas, todas con tipología correcta, campos completos y DOIs verificados.

Nota didáctica: La gracia del ejercicio está en comparar este resultado esperado con el que produzca cada participante al ejecutar el prompt. Las diferencias revelan los puntos donde el modelo duda, las decisiones que no son unívocas (¿conservar Baker 2011 o 2018?) y los límites de la auditoría automatizada.

📄 Ver .bib corregido completo

referencias_limpias.bib · 14 entradas
%% ═══════════════════════════════════════════════════════════════ %% FICHERO CORREGIDO — Auditoría bibliográfica con LLM %% 14 entradas únicas · 0 errores · Translation Studies %% ═══════════════════════════════════════════════════════════════ %% ── Baker: corregido tipo @article → @book (3ª edición) ────── %% ── Eliminado duplicado baker2011other (edición anterior) ──── @book{baker2018routledge, author = {Baker, Mona}, title = {In Other Words: A Coursebook on Translation}, publisher = {Routledge}, year = {2018}, edition = {3rd}, address = {London}, doi = {10.4324/9781315619187} } %% ── Munday: corregido tipo @article → @book ───────────────── @book{munday2016introducing, author = {Munday, Jeremy}, title = {Introducing Translation Studies: Theories and Applications}, publisher = {Routledge}, year = {2016}, edition = {4th}, address = {London} } %% ── Castilho: conservada entrada #3 (más completa) ────────── %% ── Eliminado duplicado castilho2018neural (iniciales, ────── %% año incorrecto, DOI con errata, revista abreviada) ────── @article{castilho2017machine, author = {Castilho, Sheila and Moorkens, Joss and Gaspari, Federico and Calixto, Iacer and Tinsley, John and Way, Andy}, title = {Is Neural Machine Translation the New State of the Art?}, journal = {The Prague Bulletin of Mathematical Linguistics}, year = {2017}, volume = {108}, number = {1}, pages = {109--120}, doi = {10.1515/pragl-2017-0013} } %% ── Toral & Way: sin cambios ──────────────────────────────── @article{toral2018literary, author = {Toral, Antonio and Way, Andy}, title = {What Level of Quality Can Neural Machine Translation Attain on Literary Text?}, journal = {Translation Spaces}, year = {2018}, volume = {7}, number = {1}, pages = {81--106}, doi = {10.1075/ts.00007.tor} } %% ── Nord: añadido year (1997, 1ª ed. St. Jerome) ─────────── @incollection{nord1997functionalist, author = {Nord, Christiane}, title = {Translating as a Purposeful Activity: Functionalist Approaches Explained}, booktitle = {Translation Theories Explored}, publisher = {St. Jerome}, year = {1997}, address = {Manchester} } %% ── Koponen: corregida errata título (Transltion → Translation) %% ── Corregido campo journal → booktitle ───────────────────── @inproceedings{koponen2016machine, author = {Koponen, Maarit}, title = {Machine Translation Post-Editing and Effort: An Overview}, booktitle = {Proceedings of the 19th Annual Conference of the EAMT}, year = {2016}, pages = {131--142}, address = {Riga} } %% ── Läubli: conservada entrada #9 (con DOI y diéresis) ────── %% ── Eliminado duplicado laubli2018parity (sin diéresis, ───── %% sin DOI, revista abreviada, título sin guion) ─────────── @article{laubli2018machine, author = {L{\"a}ubli, Samuel and Sennrich, Rico and Volk, Martin}, title = {Has Machine Translation Achieved Human Parity? {A} Case for Document-Level Evaluation}, journal = {Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing}, year = {2018}, pages = {4791--4796}, doi = {10.18653/v1/D18-1512} } %% ── Hurtado Albir: corregido tipo @article → @book ───────── %% ── Corregido campo journal → publisher ───────────────────── @book{hurtado2001traduccion, author = {Hurtado Albir, Amparo}, title = {Traducci{\'o}n y Traductolog{\'i}a: Introducci{\'o}n a la traductolog{\'i}a}, publisher = {C{\'a}tedra}, year = {2001}, address = {Madrid} } %% ── Moorkens: sin cambios ─────────────────────────────────── @article{moorkens2018translators, author = {Moorkens, Joss}, title = {What to Expect from Neural Machine Translation: A Practical In-Class Translation Evaluation Exercise}, journal = {The Interpreter and Translator Trainer}, year = {2018}, volume = {12}, number = {4}, pages = {375--387}, doi = {10.1080/1750399X.2018.1501639} } %% ── Venuti: añadido publisher faltante ────────────────────── @book{venuti2012translation, author = {Venuti, Lawrence}, title = {The Translation Studies Reader}, publisher = {Routledge}, year = {2012}, edition = {3rd}, address = {London} } %% ── Guerberof & Toral: sin cambios ───────────────────────── @article{guerberof2022creativity, author = {Guerberof-Arenas, Ana and Toral, Antonio}, title = {Creativity in Neural Machine Translation: A Preliminary Analysis}, journal = {Translation Spaces}, year = {2022}, volume = {11}, number = {1}, pages = {1--25}, doi = {10.1075/ts.21025.gue} } %% ── Kenny: DOI marcado como sospechoso ────────────────────── @article{kenny2020machine, author = {Kenny, Dorothy}, title = {Machine Translation for Everyone: Empowering Users in the Age of Artificial Intelligence}, journal = {Translation Studies}, year = {2020}, volume = {1}, pages = {1--16}, doi = {VERIFICAR: 10.1080/14781700.2019.9999999} } %% ── O'Brien: expandido nombre de revista, pages pendiente ─── @article{obrien2012towards, author = {O'Brien, Sharon}, title = {Towards a Dynamic Quality Evaluation Model for Translation}, journal = {The Journal of Specialised Translation}, year = {2012}, volume = {17}, pages = {VERIFICAR} } %% ── Daems et al.: sin cambios ─────────────────────────────── @article{daems2017translation, author = {Daems, Joke and Vandepitte, Sonia and Hartsuiker, Robert and Desmet, Lieve}, title = {Translation Methods and Experience: A Comparative Analysis of Human Translation and Post-Editing}, journal = {Perspectives}, year = {2017}, volume = {25}, number = {3}, pages = {452--473}, doi = {10.1080/0907676X.2016.1246377} }
4

Alternativas: RStudio y Python

Para quienes prefieran un enfoque programático, aquí van los scripts equivalentes que automatizan la detección de campos faltantes y duplicados por DOI. La auditoría semántica (duplicados con variantes de nombre, erratas en títulos) sigue requiriendo el LLM.

🟦 Script R · Auditoría básica de .bib

auditoria_bib.R
# ── Auditoría bibliográfica básica con R ───────────────────── # install.packages("bib2df") # solo la primera vez library(bib2df) library(tidyverse) # 1. CARGA ──────────────────────────────────────────────────── # Aseguramos que el .bib termina en línea vacía (evita warning) lineas <- readLines("referencias_con_errores.bib", warn = FALSE) tmp <- tempfile(fileext = ".bib") writeLines(lineas, tmp) bib <- bib2df(tmp) cat("Total entradas:", nrow(bib), "\n\n") # 2. CAMPOS FALTANTES POR TIPO ──────────────────────────────── campos_req <- list( ARTICLE = c("AUTHOR","TITLE","JOURNAL","YEAR","VOLUME","PAGES"), BOOK = c("AUTHOR","TITLE","PUBLISHER","YEAR"), INCOLLECTION = c("AUTHOR","TITLE","BOOKTITLE","PUBLISHER","YEAR"), INPROCEEDINGS = c("AUTHOR","TITLE","BOOKTITLE","YEAR") ) # Función auxiliar: evita la colisión de \(c) con la primitiva c() detectar_faltantes <- function(fila, reqs) { tipo <- toupper(fila$CATEGORY) req <- reqs[[tipo]] if (is.null(req)) return(character(0)) faltantes <- character(0) for (campo in req) { val <- fila[[campo]] if (is.null(val) || length(val) == 0 || all(is.na(val)) || val == "") { faltantes <- c(faltantes, campo) } } faltantes } cat("=== CAMPOS FALTANTES ===\n") for (i in seq_len(nrow(bib))) { fila <- bib[i, ] falt <- detectar_faltantes(fila, campos_req) if (length(falt) > 0) { cat(sprintf(" [%s] (%s) → Faltan: %s\n", fila$BIBTEXKEY, toupper(fila$CATEGORY), paste(falt, collapse = ", "))) } } # 3. DUPLICADOS POR DOI ─────────────────────────────────────── con_doi <- bib |> filter(!is.na(DOI) & DOI != "") dups_doi <- con_doi |> group_by(DOI) |> filter(n() > 1) cat("\n=== DUPLICADOS POR DOI ===\n") if (nrow(dups_doi) > 0) { for (d in unique(dups_doi$DOI)) { claves <- dups_doi |> filter(DOI == d) |> pull(BIBTEXKEY) cat(sprintf(" DOI %s → %s\n", d, paste(claves, collapse = " / "))) } } else { cat(" Ningún duplicado exacto por DOI\n") } # 4. DUPLICADOS APROXIMADOS POR TÍTULO ──────────────────────── bib <- bib |> mutate(titulo_norm = tolower(gsub("[^a-z0-9]", "", TITLE))) dups_titulo <- bib |> group_by(titulo_norm) |> filter(n() > 1) cat("\n=== POSIBLES DUPLICADOS POR TÍTULO ===\n") if (nrow(dups_titulo) > 0) { for (t in unique(dups_titulo$titulo_norm)) { claves <- dups_titulo |> filter(titulo_norm == t) |> pull(BIBTEXKEY) cat(sprintf(" «%s...» → %s\n", substr(t, 1, 50), paste(claves, collapse = " / "))) } } # 5. ENTRADAS @article CON CAMPOS DE @book ──────────────────── sospechosas <- bib |> filter(CATEGORY == "article" & !is.na(PUBLISHER) & PUBLISHER != "") cat("\n=== @article CON CAMPO publisher (¿debería ser @book?) ===\n") if (nrow(sospechosas) > 0) { for (i in seq_len(nrow(sospechosas))) { cat(sprintf(" [%s] publisher = '%s'\n", sospechosas$BIBTEXKEY[i], sospechosas$PUBLISHER[i])) } } else { cat(" Ninguna detectada\n") } cat("\n=== Auditoría completada ===\n") cat("Para detección semántica avanzada, usa el prompt de Claude.\n")

🐍 Script Python · Auditoría básica de .bib

auditoria_bib.py
# ── Auditoría bibliográfica básica con Python ──────────────── # pip install bibtexparser import bibtexparser import re from collections import Counter # 1. CARGA ──────────────────────────────────────────────────── with open("referencias_con_errores.bib", encoding="utf-8") as f: bib_db = bibtexparser.load(f) entradas = bib_db.entries print(f"Total entradas: {len(entradas)}\n") # 2. CAMPOS FALTANTES POR TIPO ──────────────────────────────── CAMPOS_REQ = { "article": ["author","title","journal","year","volume","pages"], "book": ["author","title","publisher","year"], "incollection": ["author","title","booktitle","publisher","year"], "inproceedings": ["author","title","booktitle","year"], } print("=== CAMPOS FALTANTES ===") for e in entradas: tipo = e.get("ENTRYTYPE", "").lower() req = CAMPOS_REQ.get(tipo, []) faltantes = [c for c in req if not e.get(c, "").strip()] if faltantes: print(f" [{e['ID']}] ({tipo}) → Faltan: {', '.join(faltantes)}") # 3. DUPLICADOS POR DOI ─────────────────────────────────────── dois = [(e["ID"], e["doi"]) for e in entradas if e.get("doi","").strip()] doi_count = Counter(d for _, d in dois) dups = {d for d, n in doi_count.items() if n > 1} print("\n=== DUPLICADOS POR DOI ===") for d in dups: claves = [clave for clave, doi in dois if doi == d] print(f" DOI {d} → {' / '.join(claves)}") # 4. DUPLICADOS POR TÍTULO NORMALIZADO ──────────────────────── def normalizar(t): return re.sub(r"[^a-z0-9]", "", t.lower()) titulos = [(e["ID"], normalizar(e.get("title",""))) for e in entradas] titulo_count = Counter(t for _, t in titulos) dups_t = {t for t, n in titulo_count.items() if n > 1} print("\n=== POSIBLES DUPLICADOS POR TÍTULO ===") for t in dups_t: claves = [clave for clave, titulo in titulos if titulo == t] print(f" «{t[:50]}...» → {' / '.join(claves)}") # 5. @article CON PUBLISHER (sospecha de @book) ────────────── print("\n=== @article CON publisher (¿debería ser @book?) ===") for e in entradas: if e.get("ENTRYTYPE") == "article" and e.get("publisher","").strip(): print(f" [{e['ID']}] publisher = '{e['publisher']}'") # 6. DOIs SOSPECHOSOS ───────────────────────────────────────── print("\n=== DOIs CON FORMATO SOSPECHOSO ===") for e in entradas: doi = e.get("doi", "") if doi and not re.match(r"^10\.\d{4,9}/[^\s]+$", doi): print(f" [{e['ID']}] doi = '{doi}'") print("\n=== Auditoría completada ===") print("Para detección semántica avanzada, usa el prompt de Claude.")