HTMLRewriter
Use Bun's HTMLRewriter to transform HTML documents with CSS selectors
HTMLRewriter transforms HTML documents with CSS selectors. It works with Response, string, and ArrayBuffer inputs. Bun's implementation is based on Cloudflare's lol-html.
Usage#
A common use case is rewriting URLs in HTML content:
// Replace all images with a rickroll
const rewriter = new HTMLRewriter().on("img", {
element(img) {
// Famous rickroll video thumbnail
img.setAttribute("src", "https://img.youtube.com/vi/dQw4w9WgXcQ/maxresdefault.jpg");
// Wrap the image in a link to the video
img.before('<a href="https://www.youtube.com/watch?v=dQw4w9WgXcQ" target="_blank">', {
html: true,
});
img.after("</a>", { html: true });
// Add some fun alt text
img.setAttribute("alt", "Definitely not a rickroll");
},
});
// An example HTML document
const html = `
<html>
<body>
<img src="/cat.jpg">
<img src="dog.png">
<img src="https://example.com/bird.webp">
</body>
</html>
`;
const result = rewriter.transform(html);
console.log(result);The rewriter replaces every image with a thumbnail of Rick Astley and wraps each <img> in a link, producing a diff like this:
<html>
<body>
<img src="/cat.jpg" />
<img src="dog.png" />
<img src="https://example.com/bird.webp" />
<a href="https://www.youtube.com/watch?v=dQw4w9WgXcQ" target="_blank">
<img src="https://img.youtube.com/vi/dQw4w9WgXcQ/maxresdefault.jpg" alt="Definitely not a rickroll" />
</a>
<a href="https://www.youtube.com/watch?v=dQw4w9WgXcQ" target="_blank">
<img src="https://img.youtube.com/vi/dQw4w9WgXcQ/maxresdefault.jpg" alt="Definitely not a rickroll" />
</a>
<a href="https://www.youtube.com/watch?v=dQw4w9WgXcQ" target="_blank">
<img src="https://img.youtube.com/vi/dQw4w9WgXcQ/maxresdefault.jpg" alt="Definitely not a rickroll" />
</a>
</body>
</html>Clicking any image now leads to a very famous video.
Input types#
HTMLRewriter can transform HTML from several input types:
// From Response
rewriter.transform(new Response("<div>content</div>"));
// From string
rewriter.transform("<div>content</div>");
// From ArrayBuffer
rewriter.transform(new TextEncoder().encode("<div>content</div>").buffer);
// From Blob (wrap in a Response)
rewriter.transform(new Response(new Blob(["<div>content</div>"])));
// From File (wrap in a Response)
rewriter.transform(new Response(Bun.file("index.html")));The Cloudflare Workers implementation of HTMLRewriter only supports Response objects.
Element Handlers#
The on(selector, handlers) method registers handlers for HTML elements that match a CSS selector. The handlers run for each matching element during parsing:
rewriter.on("div.content", {
// Handle elements
element(element) {
element.setAttribute("class", "new-content");
element.append("<p>New content</p>", { html: true });
},
// Handle text nodes
text(text) {
text.replace("new text");
},
// Handle comments
comments(comment) {
comment.remove();
},
});Handlers can be asynchronous and return a Promise. The transformation pauses on that element until the Promise settles, so handlers still run one at a time, in document order:
rewriter.on("div", {
async element(element) {
const fragment = await fetch("https://example.com/fragment").then(r => r.text());
element.setInnerContent(fragment, { html: true });
},
});transform(response) returns immediately. The rewrite continues in the
background, and you read the result off the returned Response. Reading it
paces the rewrite: a streamed input (a file, a fetch() response, a
ReadableStream) is pulled through only as fast as the returned body is
consumed, so a slow reader does not accumulate the whole document in memory. If
nothing reads the body, the rewrite still runs every handler to the end of the
document and buffers the output until it is read. Because the rewrite outlives
transform(), an error thrown by an async handler (or a Promise it returns that
rejects) rejects the response body instead of throwing from transform():
try {
const output = await rewriter.transform(new Response(html)).text();
} catch (error) {
console.error("a handler failed:", error);
}transform() on a string or ArrayBuffer has to return its result
synchronously, so it cannot wait for a handler that needs the event loop to turn
(a timer, I/O, a fetch). Such a handler makes transform() throw a
TypeError, and the rewrite fails without running any further handlers:
new HTMLRewriter()
.on("div", {
async element(element) {
await Bun.sleep(1000); // needs a timer
},
})
.transform("<div></div>");
// TypeError: HTMLRewriter.transform() cannot synchronously return a string
// because a content handler returned a Promise that did not resolve within a
// microtask. Pass a Response instead and await its bodyA handler whose Promise settles within a microtask checkpoint still works with
transform(string). Anything that does not need the event loop qualifies,
including process.nextTick and already-resolved Promises. Pass a Response
whenever a handler might await real work.
CSS Selector Support#
The on() method supports a wide range of CSS selectors:
// Tag selectors
rewriter.on("p", handler);
// Class selectors
rewriter.on("p.red", handler);
// ID selectors
rewriter.on("h1#header", handler);
// Attribute selectors
rewriter.on("p[data-test]", handler); // Has attribute
rewriter.on('p[data-test="one"]', handler); // Exact match
rewriter.on('p[data-test="one" i]', handler); // Case-insensitive
rewriter.on('p[data-test="one" s]', handler); // Case-sensitive
rewriter.on('p[data-test~="two"]', handler); // Word match
rewriter.on('p[data-test^="a"]', handler); // Starts with
rewriter.on('p[data-test$="1"]', handler); // Ends with
rewriter.on('p[data-test*="b"]', handler); // Contains
rewriter.on('p[data-test|="a"]', handler); // Dash-separated
// Combinators
rewriter.on("div span", handler); // Descendant
rewriter.on("div > span", handler); // Direct child
// Pseudo-classes
rewriter.on("p:nth-child(2)", handler);
rewriter.on("p:first-child", handler);
rewriter.on("p:nth-of-type(2)", handler);
rewriter.on("p:first-of-type", handler);
rewriter.on("p:not(:first-child)", handler);
// Universal selector
rewriter.on("*", handler);Element Operations#
All element modification methods return the element instance, so you can chain calls:
rewriter.on("div", {
element(el) {
// Attributes
el.setAttribute("class", "new-class").setAttribute("data-id", "123");
const classAttr = el.getAttribute("class"); // "new-class"
const hasId = el.hasAttribute("id"); // boolean
el.removeAttribute("class");
// Content manipulation
el.setInnerContent("New content"); // Escapes HTML by default
el.setInnerContent("<p>HTML content</p>", { html: true }); // Parses HTML
el.setInnerContent(""); // Clear content
// Position manipulation
el.before("Content before").after("Content after").prepend("First child").append("Last child");
// HTML content insertion
el.before("<span>before</span>", { html: true })
.after("<span>after</span>", { html: true })
.prepend("<span>first</span>", { html: true })
.append("<span>last</span>", { html: true });
// Removal
el.remove(); // Remove element and contents
el.removeAndKeepContent(); // Remove only the element tags
// Properties
console.log(el.tagName); // Lowercase tag name
console.log(el.namespaceURI); // Element's namespace URI
console.log(el.selfClosing); // Whether element is self-closing (e.g. <div />)
console.log(el.canHaveContent); // Whether element can contain content (false for void elements like <br>)
console.log(el.removed); // Whether element was removed
// Attributes iteration
for (const [name, value] of el.attributes) {
console.log(name, value);
}
// End tag handling
el.onEndTag(endTag => {
endTag.before("Before end tag");
endTag.after("After end tag");
endTag.remove(); // Remove the end tag
console.log(endTag.name); // Tag name in lowercase
});
},
});Text Operations#
Text chunks represent portions of text content and report their position in the text node:
rewriter.on("p", {
text(text) {
// Content
console.log(text.text); // Text content
console.log(text.lastInTextNode); // Whether this is the last chunk
console.log(text.removed); // Whether text was removed
// Manipulation
text.before("Before text").after("After text").replace("New text").remove();
// HTML content insertion
text
.before("<span>before</span>", { html: true })
.after("<span>after</span>", { html: true })
.replace("<span>replace</span>", { html: true });
},
});Comment Operations#
Comments support similar methods to text nodes:
rewriter.on("*", {
comments(comment) {
// Content
console.log(comment.text); // Comment text
comment.text = "New comment text"; // Set comment text
console.log(comment.removed); // Whether comment was removed
// Manipulation
comment.before("Before comment").after("After comment").replace("New comment").remove();
// HTML content insertion
comment
.before("<span>before</span>", { html: true })
.after("<span>after</span>", { html: true })
.replace("<span>replace</span>", { html: true });
},
});Document Handlers#
The onDocument(handlers) method registers handlers for events at the document level rather than within specific elements:
rewriter.onDocument({
// Handle doctype
doctype(doctype) {
console.log(doctype.name); // "html"
console.log(doctype.publicId); // public identifier if present
console.log(doctype.systemId); // system identifier if present
},
// Handle text nodes
text(text) {
console.log(text.text);
},
// Handle comments
comments(comment) {
console.log(comment.text);
},
// Handle document end
end(end) {
end.append("<!-- Footer -->", { html: true });
},
});Response Handling#
When transforming a Response, HTMLRewriter:
- Preserves the status code, headers, and other response properties
- Transforms the body while maintaining streaming capabilities
- Handles content-encoding (like gzip) automatically
- Marks the original response body as used after transformation
- Clones headers to the new response
Error Handling#
The overload you called decides which channel an error takes. Timing never
does. transform() itself throws for:
- Invalid selector syntax in the
on()method - Invalid input types (for example, passing a Symbol)
- Body already used errors, and input bodies that have already failed or aborted
- Anything a content handler raises on a
string/ArrayBufferinput, since those have to produce their result beforetransform()returns. The same goes for a handler that needs the event loop (see Element Handlers)
try {
const result = rewriter.transform("<div></div>");
} catch (error) {
console.error("HTMLRewriter error:", error);
}For a Response input, transform() returns before the rewrite finishes, so
everything the rewrite discovers surfaces on the output body instead:
- An error thrown by a content handler, or a rejected Promise one returned
- Malformed or truncated input
- Stream errors reading the input body
- Memory allocation failures
try {
const output = await rewriter.transform(new Response(html)).text();
} catch (error) {
console.error("the rewrite failed:", error);
}If a handler creates a Promise but neither returns nor awaits it, a rejection
from that Promise reaches neither channel. Like any detached rejection, it goes
to the process-global unhandledRejection path. Earlier versions of Bun could
surface it from transform() itself.
See also#
You can also read the Cloudflare documentation, which this API is intended to be compatible with.