oltarasenko
Crawly - A high-level web crawling & scraping framework for Elixir
Dear Elixir community,
After a year of development, bug fixes, and improvements, we are proudly ready to share the release of Crawly 0.10.0 here with you.
Check out the source code here: GitHub - elixir-crawly/crawly: Crawly, a high-level web crawling & scraping framework for Elixir. · GitHub
Check out our experimental, pre-alpha visual UI here: http://18.216.221.122/ and play with jobs scheduling there.
We have dedicated a lot of time and knowledge to build a fast and feature-rich web scraping framework. To be absolutely honest, I have to say that I took a lot of ideas from other popular web scraping framework (Scrapy, python), as I have previously worked with the Scrapy core team.
We have some reported production usages of Crawly (some of them with really long-running crawls), however, we still have to approach the stable version (hopefully it will happen in a couple of releases).
To describe how Crawly is different from other known elixir scraping frameworks, I will list crawly’s features which I believe make it outstanding:
- Documentation - we have spent an enormous amount of time and effort to build great and clear versioned documentation!
- Rate limiting
- Robots.txt support
- Requests and Items validators
- Automatic duplications filtering
- Automatic cookies management (allows to bypass login pages and cookie-based regional filtering)
- Browser rendering (with the help of Splash)
- Retries support
- Proxies support
- HTTP API
- Visual jobs management dashboard which allows operating multiple Crawly nodes at the same time (experimental): see it deployed on demo ec2 micro instance: http://18.216.221.122/
We hope it will be useful for you!
If you have a suggestion or a production use case you’re happy for us to share, please get in touch.
First Post!
dgreiss
Most Liked
oltarasenko
Hey people,
I’ve got a bit of time to work on Crawly recently, and as a result have made a new release and a new article about it ![]()
Hopefully you will find it useful.
oltarasenko
As a part of the project development, I have decided to create a short cookbook of scraping recipes, if you’re using Crawly or Scraping in general, these articles might be useful for you (I am including medium friend links, so everyone can read them):
https://oltarasenko.medium.com/web-scraping-with-elixir-and-crawly-browser-rendering-afcaacf954e8?sk=345c1777d793aa99828986010cc3b077
oltarasenko
I can also refer my talk from November 2019 explaining how scraping can be used in general: https://www.youtube.com/watch?v=ovSQGlkakAQ
Last Post!
RicoTrevisan
Thanks, indeed that is working as you mentioned. I was in dev starting / stopping the spider manually and never gave it a chance to stop properly.
In any case, after more consideration, I’ve decided to keep it off.
I’m using Crawly in a Phoenix application. I see Crawly automatically looks for spiders in ./spiders. I normally put all my backend modules in ./lib/my_app. Is there a way to change the default folder from ./spiders to ./lib/my_app/spiders?
Thanks again for Crawly.
I’m hoping to contribute back to it – as soon as I figure out how to better display the docs graphs in both Hex and in the IDE.
Popular in Announcing
Other popular topics
Categories:
Sub Categories:
Forums
Popular Tags
- #ecto
- #liveview
- #troubleshooting
- #learning-elixir
- #deployment
- #library
- #erlang
- #testing
- #genserver
- #mix
- #absinthe
- #remote-other
- #otp
- #plug
- #how-to-question
- #macros
- #postgres
- #channels
- #elixirconf
- #exunit
- #discussion
- #code-sync
- #javascript
- #podcasts
- #onsite
- #dialyzer
- #docker
- #authentication
- #umbrella
- #full-time-contract
- #podcasts-by-brainlid
- #ecto-query
- #elixir-ls
- #phoenix_html
- #iex
- #blog-post
- #graphql
- #genstage
- #ai
- #websockets
- #supervisor
- #elixirconf-us
- #advent-of-code
- #distillery
- #processes
- #api
- #forms
- #metaprogramming
- #security
- #hex









