Security Considerations
xlsx-format is designed for modern XLSX reading and writing in Node.js and browsers. Spreadsheet files are still untrusted input: treat uploads, email attachments, partner feeds, and user-edited workbooks as data from an attacker-controlled boundary.
What xlsx-format does not execute
xlsx-format parses workbook structures and cell data. It does not execute formulas, VBA macros, external entities, embedded scripts, or workbook event handlers.
That avoids common execution classes such as XXE and macro execution, but it does not remove resource-exhaustion or unsafe-export risks. A malicious workbook can still try to consume excessive CPU or memory with large ZIP payloads, large XML parts, huge declared sheet ranges, or many repeated strings.
Reading untrusted files
When reading files from users or external systems:
- apply upload size limits before calling
read(), - run parsing away from latency-sensitive request paths when possible,
- use process, worker, or request timeouts in server environments,
- prefer narrow conversion work instead of immediately materializing every sheet into JSON,
- use
sheetRowswhen only a preview or header sample is needed, - tune ZIP and XML parser limits for public upload paths,
- keep xlsx-format updated for ZIP, XML, and export hardening fixes.
import { read, sheetToJson } from "xlsx-format";
const workbook = await read(fileBytes, {
sheetRows: 1000,
maxXmlPartBytes: 32 * 1024 * 1024,
maxXmlTags: 500_000,
maxXmlNestingDepth: 128,
maxSharedStringItems: 100_000,
maxWorksheetRows: 100_000,
maxWorksheetCells: 1_000_000,
});
const firstSheet = workbook.Sheets[workbook.SheetNames[0]];
const previewRows = sheetToJson(firstSheet);The default ZIP and XML limits are intentionally high enough for ordinary workbooks. Lower them when a service receives arbitrary uploads and does not need to parse very large sheets.
maxWorksheetRows limits row elements or text records scanned; it does not reject a valid cell solely because its row number is high. sheetRows is the separate position-based preview limit. maxWorksheetCells is a cumulative per-sheet work budget: it counts explicit cells and work generated from column definitions, hyperlink ranges, and HTML spans. XLSX and text imports default to 10,000,000 cell work units.
Worksheet exporters, including CSV, JSON, HTML, formula lists, and XLSX writing, clamp oversized declared ranges to occupied cells and then enforce a 1,000,000-position default. XLSX writing also charges row and column metadata entries, and HTML writing charges merged-range metadata entries. Pass maxWorksheetCells to an exporter or write() when a trusted worksheet intentionally requires a larger range. An occupied cell beyond the configured budget causes LIMIT_EXCEEDED; it is never silently discarded.
Password-protected workbooks
xlsx-format does not read or write encrypted or password-protected workbooks. Passing a non-empty password option to read() or write() throws an XlsxError with code UNSUPPORTED before parsing or serialization starts. Omit the option for ordinary unencrypted workbooks.
Handling errors
xlsx-format throws XlsxError, a subclass of Error, for deterministic library failures. Existing catch blocks keep working, but new integrations should branch on error.code instead of matching message text.
import { read, XlsxError } from "xlsx-format";
try {
const workbook = await read(fileBytes);
} catch (error) {
if (error instanceof XlsxError) {
if (error.code === "MALFORMED" || error.code === "CRC_MISMATCH") {
// Reject corrupt uploads.
}
if (error.code === "LIMIT_EXCEEDED") {
// Keep public upload limits low, or retry trusted files with higher ReadOptions limits.
}
}
throw error;
}Stable error codes:
| Code | Meaning |
|---|---|
INVALID_ARGUMENT | Caller-supplied options or API inputs are invalid. |
MALFORMED | The workbook, ZIP, XML, or format structure is corrupt. |
LIMIT_EXCEEDED | A configured ZIP, XML, worksheet, or shared-string limit failed. |
UNSUPPORTED | The input uses a valid but unsupported format or feature. |
CRC_MISMATCH | ZIP entry data failed its CRC integrity check. |
NOT_FOUND | A required workbook part or worksheet reference is missing. |
DUPLICATE | A workbook or archive structure contains a duplicate entry. |
A full XLSX read rejects missing selected worksheets and missing shared string tables declared by the package. Sheet-name-only and property-only reads do not load those parts. Optional worksheet relationships, comments, and legacy drawing annotations remain best-effort when WTF is false; set WTF: true to surface malformed optional parts. Configured limits and invalid option values always throw, regardless of WTF.
Exporting CSV
CSV files opened in spreadsheet applications can interpret text that starts with characters such as =, +, -, @, tab, or carriage return as formulas. This is commonly called CSV injection or formula injection.
By default, sheetToCsv() and CSV output from write() prefix formula-like text with a single quote. Pass escapeFormulae: false only when you need exact text fidelity and the output will not be opened as an untrusted spreadsheet.
import { sheetToCsv } from "xlsx-format";
const csv = sheetToCsv(sheet);Exporting HTML
HTML export escapes cell text and attribute values and, by default, drops hyperlink targets with unsafe schemes such as javascript:, vbscript:, and data:. Rich-text HTML cached in CellObject.h, including values supplied directly by an application, is limited to the presentation tags and styles generated by the XLSX rich-text reader. Unsupported tags are rendered as text. The header and footer options are application-supplied HTML and must contain only markup your application trusts. Pass sanitizeLinks: false only when you need exact hyperlink fidelity and will not render the exported HTML in a browser. If your application accepts arbitrary hyperlinks, validate links against your own allowed URL schemes and domains before publishing or embedding the exported HTML.
htmlToSheet() is a lightweight table importer rather than a browser DOM parser. It recognizes quoted cell attributes, ordinary <br> line breaks, and semicolon-terminated WHATWG named and numeric character references. Use a browser or standards-compliant HTML parser first when accepting arbitrary documents that depend on broader HTML error recovery.
import { sheetToHtml } from "xlsx-format";
const html = sheetToHtml(sheet);Object keys from workbook data
Spreadsheet headers and sheet names can become object keys in APIs such as sheetToJson() and workbook.Sheets. Treat keys from untrusted files as untrusted data. Avoid merging parsed row objects directly into application configuration, prototypes, or privileged domain objects.
const rows = sheetToJson<Record<string, unknown>>(sheet);
for (const row of rows) {
const safeName = String(row["Name"] ?? "");
// Copy only the fields your application expects.
}Reporting vulnerabilities
See the Security Policy for supported versions, private reporting guidance, and disclosure expectations.