Posted in

Building a Web Scraper with NodeJS

A hand holding an iPhone displaying the Siri glowing orb interface on screen with the Google logo and text "Powered by Google Gemini" at the bottom.
A look at the interface architecture of the 2026 Apple and Google partnership, bringing Gemini's frontier models to Siri.

Web scraping is a great way to extract data from a large website to obtain specific research and comparative purposes. For example, you may want to view all the details for a line of products or find statistics on sports teams or hotel rates.

While there is the option to scrape the web for data manually, it is time-consuming. If you’re wondering how to build a web scraper and worried it might be complicated, don’t be. Data is easily extracted with an automatic tool to make the task fast and easy. Building this tool requires just a few steps to establishing a search engine to find what you’re looking to query. This process can be successfully done with a NodeJS web scraping tutorial.

Preparing to Build: What You Need

There are two significant steps to building your web scraper:

  • Find the data you need through an HTML request library or use a browser without a heading.

  • Begin parsing the raw data to find the specific information you need.

Here’s what you need to start:

  • A version 8 or higher of NodeJS and npm installed onto your computer.

  • Libraries for requesting web scraping data (request-promise or Axios, Puppeteer, and Cheerio)

To set up the preliminary stage, you’ll need to download three libraries: request-promise, Puppeteer, and CheerioJS with the npm install command, followed by “save request request-promise cheerio puppeteer. Downloading time frames can vary, especially with Puppeteer, because it requires Chromium to install at the same time. Once all installations are complete, you’ll be ready to make your first request.

Step 1: Make Your First Request with a Specific Search

Once you have all the libraries installed, you’re ready to make your first query request. Open a new text file to write a function to receive the HTML from the desired website. You’ll need to determine which specific data you are searching for and write a function to acquire the HTML for a list or website.

For example, if you are looking for hotel or resort listings, you may request a vacation website such as Kayak or Expedia. Name the file based on the search criteria, i.e., hotelScraper.js. Wikipedia is a popular choice for finding large amounts of lists or data:

const URL = HTTPS:// (name of website)

Step 2: Working Through the Raw Data

Once you receive the output or raw HTML from the website you’re looking to filter, you’ll notice a large amount of text that’s difficult to decipher. Chrome DevTools is the next tool to use to filter through the HTML of the web page. This process is done through Google Chrome by selecting the element to scrape using a right-click. For example, if you want to “scrape” hotels, click on this element to receive the links to individual pages.

The results will appear in a separate pane with DevTools to show you everything related to the page’s HTML source. When reviewing the data, look for the “big” tag indicating a hyperlink inside. Parsing the raw data is the next stage of the process, and this is done with Cheerio.js.

Step 3: Parsing the Data with Cheerio.js

Cheerio.js is the tool used to parse the HTML on the initial list, showing the website page’s links. This command is made as follows:

const rp = require ( request-promise’)

const $ = require ( cheerio’)

const URL = HTTPS:// (website address details)

The results or output lists all the elements without any “big” tags on the page. This list is more detailed, with further links to specific pages for each element included. This next query provides details with all the individual websites related to each element from the previous list.

For instance, if you searched for hotel chains located in upstate New York, you can parse by creating a new file, i.e., hotelParse.js, to indicate which specific data you want to exact from each source:

const rp = require( request-promise’)

const URL = HTTPS:// (specific website from element “ i.e. hotel website)

Step 4: Defining Data for Further Parsing

This command specifies the element’s link with each hotel’s information, such as the price range, room sizes, location, etc. To find this information, Chrome DevTools needs to find the syntax of the code for parsing. If you are looking for hotel addresses, for example, you can extract this by using Cheerio.js:

const rp = require( request-promise’)

const $ = require( cheerio’)

const URL = HTTPS:// (specific website from element “ i.e. hotel website)

Then, add the function to extract the specific data for the hotel address, as an example, and labeling as “firstHeading”:

rp (URL)

.then(function(HTML)

console.log($( .firstHeading’, HTML).text())

console.log($( .(specific details “ i.e. hotel address or location), HTML).text())

Step 5: Compiling the Extracted Data for the Module

When you move the extracted data to a module, the next command will export and compile the information requested.

const rp = require( request-promise’)

const $ = require( cheerio’)

const (created file “ i.e. hotelParse) = function(URL)

return rp(URL)

.then(function(HTML)

return

name: $( .firstHeading’, HTML).text()

(hotel criteria “ address or location)

The above commands may vary depending on the specific data you wish to extract.

Step 6: Applying the Net Scraper File and Parse Module

Next, the two files created for the web scraping process, (i.e. hotelScraper.js and hotelParse.js). The initial file and module are combined in the next command:

const rp = require( request-promise’)

const $ = require( cheerio’)

const (hotel address)Parse = require( ./(hotel address)Parse’)

const url = https:// (website list)

rp(url)

.then(function(html)

//success!

You’ll find there are more details following this result, based on individual scraping criteria. The results, or output, should produce a detailed list of all the specific details of your web scraping search with each hotel and the address or location, or the information requested.

The Practical and Easy Way to Extract Data with NodeJS and Puppeteer

When you use the request-promise module along with the Cheerio.js, you’ll find it easy to scrape data from most websites online. If a website uses JavaScript for its content, this may cause issues with extracting data using HTTP-based request libraries. By using the Puppeteer module, allows you to run scraping on JavaScript and produce results:

const puppeteer = require( puppeteer’)

const URL = HTTPS:// (website)

Puppeteer

.launch()

.then(function(browser)

return browser .newPage()

.then(function(page)

return page.goto(URL).then(function()

return page.content()

The above section may vary and require additional information based on your search or the data you are scraping to find. Building your web scraper may take some practice and changes to the commands used based on the websites you want to query.

Building Your Web Scraper Versus Using A Pre-Built Module

Pre-built web scrapers are a far easier option if you don’t have the time or don’t feel comfortable building your own.

Building your scraper can take time to learn specific tools and the options available. You’ll find variations on how to scrape for data with a NodeJS server by using various libraries and modules to acquire the desired results. There are many great resources and tutorials available to determine which tools are best for your data extraction. We also recommend using a build-and-deploy cloud-based PAAS to get your scraping tool up and running as quickly as possible.

There are many other options to consider for sifting through data on the web, from the cloud to local storage resources. The level of the user-friendly interface varies widely from one tool to another. Some platforms are much easier to use than others, with some systems requiring more complex commands than others. Some tools offer tutorial-style tips and help from one step to the next until you find the results you need. 

Jonathan specializes in SEO - his work spans multiple industries, from SaaS to Sales, and he strives for excellence. When he isn't working, Jonathan enjoys adventure and the outdoors, a real nature enthusiast. 

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.