gpatankar21

gpatankar21

I am required to parse a resume in pdf format to extract fields like phone-number github-url linkedIn-url etc, is there any way to parse the pdf to extract this data from the pdf.

Showing Posts 1 to 10

NobbZ

NobbZ

I tend to say no, unless you have a proper FORM, but still it won’t be easy then AFAIK, but I have not used any tools that would do so, as we read and process our resumes manually at my company.

What would make it hard, is that in a PDF not necessarily a single letter of the text you read as a human has to be saved in the text. In theory, it could be drawn as a single large vector graphic.

Its rarely done due to the cost in size though, and embedding fonts and text is quite common. Still, tabular views are not saved like that necessarily.

They could be saved as a single free positioned box per cell, without any possibility to read programatically which row and column this cell belongs to, but easy recognizable as a human.

Just read the resume, or if you have a license for a good OCR software, use that. Or require resumes to be handed in in a better machine readable format.

outlog

outlog

It’s difficult if the pdf is not systematically created by some system (say bank statement pdfs from your bank)

see Lib for pdf processing - #2 by AstonJ - and experiment with parsing the output from pdftotext - eg linkedin uri and emails should be easy - something like phone number, addresses will be more difficult.

tme_317

tme_317

I use Ghostscript to convert the PDF to a txt file then attempt to parse it.

System.cmd("gs", ["-sDEVICE=txtwrite", "-o#{txt_path}", pdf_path])

Parsing this resulting text file can be difficult if the PDFs are not of consistent format but at least you have text to work with.

gpatankar21

gpatankar21 OP

I used pdftotext to convert the content of pdf but it does not worlb for large sized pdf’s
It works when the pdf size is about 5 kb or less. Otherwise the content is empty in pdf_as_text.

{pdf_as_text,_} = System.cmd("pdftotext", ~w[#{attrs["attachment"].filename} -], cd: "/home/gagan/aviahire-web/uploads/gagan/applicant/3/attachments/thumb")
      IO.puts("++++++++")
      IO.puts(pdf_as_text)
      IO.puts("++++++++")
NobbZ

NobbZ

What is the exit code? You really shouldn’t ignore it.

And what do you see when you try to rectify the bigger PDFs manually from the terminal?

gpatankar21

gpatankar21 OP

If I try to convert bigger pdfs then, I get an empty string in the pdf_as_text variable.
Is there any way I can convert bigger pdfs into text and then extract contents like url and phone number.

peerreynders

peerreynders

I get an empty string in the pdf_as_text variable.

That isn’t the issue. You are using:

{pdf_as_text,_} = System.cmd(...)

You are ignoring the exit_status from System.cmd/3

You have already observed that you are getting an empty string in pdf_as_text. So use instead:

{pdf_as_text, exit_status} = System.cmd(...)

and find out what the value of exit_status is - its value may give you some hint as to why pdf_as_text is empty.

For example this lists:

0 No error.
1 Error opening a PDF file.
2 Error opening an output file.
3 Error related to PDF permissions.
99 Other error.

NobbZ

NobbZ

Besides of what @peerreynders says, please check using on a terminal, as there you might see output on stderr which is ignored by System.cmd/3 (unless combined with stdout using an option to the call).

gpatankar21

gpatankar21 OP

what is other error?
I am getting 99 as output.
How to resolve that error?

NobbZ

NobbZ

Have you tried to invoke the command from the terminal? Does it produce any output?

Where Next? Top

Trending in Questions Top

RSP87
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
nseaSeb
Hello, I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
kpanic
Hi everyone, I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding. I sta...
New
brecabral
Documentation While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
velrest
So my question is quite simple and i have found no conclusive answer on forum, google or AI. Should we use :erlang.float for Integer to ...
New
asweet-confluent
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
apz
I’m new to elixir and just tried to install the elixirLS extension for VScode(ium) and it is throwing some errors that I would like help ...
New

Other Trending Topics Top

GenericJam
Edit: 2026 May 15 - This post is archived. Mob is alive!! Main docs: mob v0.7.11 — Documentation A bit of explanation for the slightly c...
New
JesseHerrick
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
mudasobwa
I am happy to introduce the very α version of the new programming language compiled to BEAM. Welcome Cure. It has literally three kille...
New
garrison
Hobbes is a low-level distributed database for the Elixir programming language. Hobbes provides a simple, safe, and scalable storage lay...
New
marciok
Hi there! We created Gust: A task orchestrator inspired by Airflow. For those who have never heard about Aiflow, it’s a Python-based wor...
New
mhanberg
Hi everyone! The first release candidate for the Expert language server project is now available! We’ve published a press release detai...
New

We're in Beta

About us Mission Statement

Options

Thread Display Mode




Thread Preview

Skip Thread Previews