gpatankar21
I am required to parse a resume in pdf format to extract fields like phone-number github-url linkedIn-url etc, is there any way to parse the pdf to extract this data from the pdf.
Trending in Questions
I’m working on a project that simulates the bumbl example in the programming phoenix book. It acts almost like an email client. We have a...
New
Hello,
I know there is an approach for handling lists that allows for optimized traversal, but I can’t recall the specific method (somet...
New
Hi everyone,
I am toying with the idea of building a “match maker” for giving personal help to people that wants to start coding.
I sta...
New
Documentation
While reading the Scoped Routes section, I noticed that the documentation currently refers to a problem without explainin...
New
So my question is quite simple and i have found no conclusive answer on forum, google or AI.
Should we use :erlang.float for Integer to ...
New
I recently noticed that Elixir’s Logger defaults its primary log level to :debug when no :logger, :level application configuration is pre...
New
I’m new to elixir and just tried to install the elixirLS extension for VScode(ium) and it is throwing some errors that I would like help ...
New
Other Trending Topics
Edit: 2026 May 15 - This post is archived.
Mob is alive!!
Main docs: mob v0.7.11 — Documentation
A bit of explanation for the slightly c...
New
Hey, I’m Jesse and I’m the main contributor behind Dexter, a full-featured, lightning-fast Elixir LSP optimized for large codebases. It s...
New
I am happy to introduce the very α version of the new programming language compiled to BEAM.
Welcome Cure.
It has literally three kille...
New
Hobbes is a low-level distributed database for the Elixir programming language.
Hobbes provides a simple, safe, and scalable storage lay...
New
Hi there! We created Gust: A task orchestrator inspired by Airflow.
For those who have never heard about Aiflow, it’s a Python-based wor...
New
Hi everyone!
The first release candidate for the Expert language server project is now available!
We’ve published a press release detai...
New
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #library
- #deployment
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #elixirconf
- #channels
- #exunit
- #discussion
- #code-sync
- #podcasts
- #javascript
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #ai
- #elixirconf-us
- #blog-post
- #elixir-ls
- #phoenix_html
- #iex
- #graphql
- #genstage
- #websockets
- #supervisor
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #elixirconf-eu
- #metaprogramming
- #hex










Showing Posts 1 to 10- Show Best Posts
- Show All (oldest first)
- Show All (newest first)
NobbZ
I tend to say no, unless you have a proper FORM, but still it won’t be easy then AFAIK, but I have not used any tools that would do so, as we read and process our resumes manually at my company.
What would make it hard, is that in a PDF not necessarily a single letter of the text you read as a human has to be saved in the text. In theory, it could be drawn as a single large vector graphic.
Its rarely done due to the cost in size though, and embedding fonts and text is quite common. Still, tabular views are not saved like that necessarily.
They could be saved as a single free positioned box per cell, without any possibility to read programatically which row and column this cell belongs to, but easy recognizable as a human.
Just read the resume, or if you have a license for a good OCR software, use that. Or require resumes to be handed in in a better machine readable format.
outlog
It’s difficult if the pdf is not systematically created by some system (say bank statement pdfs from your bank)
see Lib for pdf processing - #2 by AstonJ - and experiment with parsing the output from pdftotext - eg linkedin uri and emails should be easy - something like phone number, addresses will be more difficult.
tme_317
I use Ghostscript to convert the PDF to a txt file then attempt to parse it.
Parsing this resulting text file can be difficult if the PDFs are not of consistent format but at least you have text to work with.
gpatankar21
I used pdftotext to convert the content of pdf but it does not worlb for large sized pdf’s
It works when the pdf size is about 5 kb or less. Otherwise the content is empty in pdf_as_text.
NobbZ
What is the exit code? You really shouldn’t ignore it.
And what do you see when you try to rectify the bigger PDFs manually from the terminal?
gpatankar21
If I try to convert bigger pdfs then, I get an empty string in the pdf_as_text variable.
Is there any way I can convert bigger pdfs into text and then extract contents like url and phone number.
peerreynders
That isn’t the issue. You are using:
You are ignoring the
exit_statusfromSystem.cmd/3You have already observed that you are getting an empty string in
pdf_as_text. So use instead:and find out what the value of
exit_statusis - its value may give you some hint as to whypdf_as_textis empty.For example this lists:
NobbZ
Besides of what @peerreynders says, please check using on a terminal, as there you might see output on
stderrwhich is ignored bySystem.cmd/3(unless combined withstdoutusing an option to the call).gpatankar21
what is other error?
I am getting 99 as output.
How to resolve that error?
NobbZ
Have you tried to invoke the command from the terminal? Does it produce any output?