Usage
You can use any folder to store the PDFs to be processed and any other to extract the results. They don’t have to be specifically named paper, data, or results; you just need to specify them when running the commands.
Using GROBID for XML Extraction¶
To extract structured XML data from PDFs using GROBID, follow these steps:
docker run --rm -p 8070:8070 lfoppiano/grobid:latest-full
This will start the GROBID service on port 8070.
- Process PDFs with GROBID Once the GROBID server is running, you can extract XML from a folder of PDFs using the following command:
curl -F input=@<path_to_pdf> "http://localhost:8070/api/processFulltextDocument" -o <output_xml>
Alternatively, for batch processing of all PDFs in a directory:
for file in <pdf_folder>/*.pdf; do
curl -F input=@$file "http://localhost:8070/api/processFulltextDocument" -o "<output_folder>/$(basename "$file" .pdf).xml"
done
Generate Keyword Cloud¶
Extracts keywords from abstracts in XML files and creates a word cloud.
Command:
python scripts/keywordCloud.py <folder_with_xmls> <output_folder>
Output: <output_folder>/keywordCloud.jpg
Chart Figures Count¶
Counts the number of figures in each XML file and generates a bar chart.
Command:
python scripts/charts.py <folder_with_xmls> <output_folder>
Output: <output_folder>/charts.jpg
Extract Links¶
Extracts links from XML files while ignoring references.
Command:
python scripts/list.py <folder_with_xmls> <output_folder>
Output: <output_folder>/links.txt