Save a Website or Web Page as HTML in Java

To save a web page as HTML in Java, load its URL into an HTMLDocument and call save(). Use HTMLSaveOptions to control linked pages and resources such as images, CSS, and JavaScript.

Why Save a Website as HTML?

Saving a web page with its related resources is useful for offline access, archiving, testing, and content analysis. The result contains the HTML document and resources allowed by the selected save settings.

Saving a website does not automatically copy every page from a domain. Aspose.HTML follows links from the loaded document according to the configured page depth and URL restrictions. It saves the loaded HTML and related web resources, not the source files of the server-side application.

How to Save a Webpage in Java

Use the following workflow to save a web page with default settings:

  1. Create an HTMLDocument from the source web page URL.
  2. Prepare the destination path for the HTML file.
  3. Call document.save(savePath) to save the document and its permitted resources.

By default, MaxHandlingDepth is 0, so linked pages are not followed. Related resources such as images, CSS, and scripts are processed, but the default ResourceUrlRestriction value SameHost excludes resources hosted on other domains.

1// Extract and save a web page with default save options in Java
2
3// Initialize an HTML document from a URL
4final HTMLDocument document = new HTMLDocument("https://docs.aspose.com/html/net/message-handlers/");
5// Prepare a path to save the downloaded file
6String savePath = "root/result.html";
7
8// Save the HTML document to the specified file
9document.save(savePath);

Configure HTMLSaveOptions for Website Saving

HTMLSaveOptions exposes ResourceHandlingOptions, which controls scripts, linked-page traversal, and URL restrictions. The following settings are central to the examples on this page:

SettingWhat it controlsDefault
JavaScriptWhether scripts are saved, ignored, discarded, or embeddedSave
MaxHandlingDepthHow deeply linked pages are followed0
PageUrlRestrictionWhich linked page URLs may be processedRootAndSubFolders
ResourceUrlRestrictionWhich image, CSS, script, and other resource URLs may be processedSameHost

Embed JavaScript in Saved HTML

The JavaScript setting determines how script resources are handled. Supported values include Save, Ignore, Discard, and Embed. The following example selects ResourceHandling.Embed to place external JavaScript content in the saved HTML output.

  1. Load the source web page into an HTMLDocument.
  2. Create HTMLSaveOptions.
  3. Set JavaScript resource handling to ResourceHandling.Embed.
  4. Save the document with the configured options.
 1// Download website using HTMLSaveOptions in Java
 2
 3// Initialize an HTML document from a URL
 4final HTMLDocument document = new HTMLDocument("https://docs.aspose.com/html/net/message-handlers/");
 5
 6// Create an HTMLSaveOptions object and set the JavaScript property
 7HTMLSaveOptions options = new HTMLSaveOptions();
 8options.getResourceHandlingOptions().setJavaScript(ResourceHandling.Embed);
 9
10// Prepare a path to save the downloaded file
11String savePath = "rootAndEmbedJs/result.html";
12
13// Save the HTML document to the specified file
14document.save(savePath, options);

Save Linked Pages with setMaxHandlingDepth()

MaxHandlingDepth controls the depth of linked pages processed from the initial document; it does not control the depth of the HTML element tree. The default value 0 saves the current document without following linked pages. Setting it to 1, as in the next example, also processes pages linked directly from the initial document. A value of -1 allows all reachable page levels permitted by the URL restrictions and should be used carefully on large sites.

 1// Save a website with limited resource depth using Java
 2
 3// Load an HTML document from a URL
 4final HTMLDocument document = new HTMLDocument("https://docs.aspose.com/html/net/message-handlers/");
 5
 6// Create an HTMLSaveOptions object and set the MaxHandlingDepth property
 7HTMLSaveOptions options = new HTMLSaveOptions();
 8options.getResourceHandlingOptions().setMaxHandlingDepth(1);
 9
10// Prepare a path for downloaded file saving
11String savePath = "rootAndAdjacent/result.html";
12
13// Save the HTML document to the specified file
14document.save(savePath, options);

Restrict Saved Pages by URL

PageUrlRestriction controls which linked page URLs may be processed. Its default value is RootAndSubFolders. The next example combines setMaxHandlingDepth(1) with setPageUrlRestriction(UrlRestriction.SameHost).

As a result, the initial document and its directly linked pages on the same host can be saved. Links beyond the configured depth or outside the permitted host are not processed.

 1// Save a website with restricted resource URLs using Java
 2
 3// Initialize an HTML document from a URL
 4final HTMLDocument document = new HTMLDocument("https://docs.aspose.com/html/net/message-handlers/");
 5
 6// Create an HTMLSaveOptions object and set MaxHandlingDepth and PageUrlRestriction properties
 7HTMLSaveOptions options = new HTMLSaveOptions();
 8options.getResourceHandlingOptions().setMaxHandlingDepth(1);
 9options.getResourceHandlingOptions().setPageUrlRestriction(UrlRestriction.SameHost);
10
11// Prepare a path to save the downloaded file
12String savePath = "rootAndManyAdjacent/result.html";
13
14// Save the HTML document to the specified file
15document.save(savePath, options);

Common Website Saving Issues

IssueLikely causeFix
Only the initial page is savedMaxHandlingDepth remains at its default value of 0.Set an appropriate positive depth when linked pages must also be processed.
An external image, stylesheet, or script is missingThe default ResourceUrlRestriction.SameHost excludes a resource hosted on another domain.Adjust the resource URL restriction only for sources the application is allowed to retrieve.
A linked page is skippedIts depth or URL does not satisfy MaxHandlingDepth and PageUrlRestriction.Review both settings together rather than changing only one of them.
Website saving processes too many pagesThe traversal depth is too broad for the site.Use a finite depth and an appropriate page URL restriction.

Related Aspose.HTML Articles

Other Platforms

Try HTML Web Applications

Aspose.HTML provides free online HTML Web Applications for converting files, extracting web data, generating HTML, and analyzing pages without installing additional software.

HTML Web Applications