Cross-Platform DOCX, RTF, HTML, and PDF Processing with NOV

Introduction

Document workflows often begin on one platform and later expand to desktop, web, and server applications. A report initially created in a Windows application may later need to be generated from a web portal, edited on a Mac, or delivered automatically as a PDF attachment. Maintaining separate document-building code for each platform makes even small changes to formatting or layout more difficult. Furthermore, operating systems provide different text processing and rendering services, which can lead to differences in document measurement, layout, and appearance.

NOV Rich Text Editor solves these problems because it is designed as a platform-independent text processing solution. This topic describes the building blocks used in the creation of NOV Rich Text Editor and how they differ from other cross-platform solutions.

The Unicode Standard

A core part of every text processing solution is the Unicode standard. Many developers mistakenly associate Unicode only with encodings such as UTF-8 and UTF-16. However, these encodings represent just one part of the standard: how character codes are represented as bytes.

Besides text encoding, the Unicode standard defines character names, scripts, and various text processing algorithms:

AlgorithmWhat it does
Word boundariesDefines default rules for finding word boundaries, useful for selection, capitalization, keyboard text navigation, and searching.
Line breakingIdentifies positions where a line may break or must break (soft and hard line breaks). This algorithm is extensively used in paragraph layout.
Grapheme boundariesGroups code points into user-perceived characters, such as a letter with combining accents or an emoji sequence. This algorithm is useful for cursor movement and deletion.
Sentence boundariesDefines default rules for identifying sentence boundaries. This algorithm is useful for sentence capitalization and searching.
Bidirectional ordering (BIDI)Determines how characters are ordered when mixing left-to-right (LTR) and right-to-left text (RTL), such as mixing English with Arabic or Hebrew.
Normalization Converts equivalent character sequences into normalized forms, such as the composite character é versus e plus a combining accent.


Nevron Open Vision internally implements all parts of the Unicode standard that are used in the controls in the suite. This ensures that all text processing is consistent at the Unicode level regardless of the operating system or .NET environment.

The OpenType and TrueType Standards

The Unicode standard defines characters, encoding, and text processing algorithms, while font standards describe how those characters are displayed. A visual representation of a character or a sequence of characters is called a glyph. Standards such as TrueType define how character code points are mapped to glyphs and how those glyphs are displayed on the screen.

TrueType and OpenType fonts can describe glyph shapes using vector outlines, which preserve image quality when scaled. These outlines are rasterized according to the font size and display resolution. Font standards also complement Unicode by defining the metrics used to position glyphs and instructions that can adjust glyph outlines to the pixel grid to improve rasterization quality. The following table summarizes some of the font information and features used by NOV Rich Text Editor and the NOV text processing stack:

Font informationDefines the font name, general metrics, PANOSE classification, etc. Useful when the font system performs font lookup as well as font fallback.
Character-to-glyph mappingThe CMAP table has numerous formats, but their general purpose is to map character codes to glyph IDs. This is useful when the character sequence is converted to a glyph run.
Glyph substitution (GSUB) The GSUB table replaces glyph sequences with precomposed ones. Common uses include supporting ligatures (predefined visual representations of a sequence of glyphs), contextual forms (used in Arabic to model initial, medial, final, and isolated forms), and script-specific letter shapes.
Font and glyph metricsAdvance widths, bearings, ascender and descender metrics. These are useful when the glyphs are positioned relative to each other for measurement and display purposes.
Glyph outlinesOutline data describes the shapes of letters and symbols so they can be scaled for different font sizes and output resolutions.
Hinting TrueType instructions can adjust outlines to the pixel grid during rasterization, helping maintain legibility even at small sizes. NOV implements its own proprietary hinter with best-in-class performance due to optimizations that are not found in open source hinters.
Glyph positioning (GPOS)Adjusts glyph placement and spacing, including kerning and the attachment of combining marks to base glyphs.


Nevron Open Vision contains a completely managed implementation of the TTF/OTF standard that is coupled with our managed Unicode implementation for increased performance. The fact that the Nevron Rich Text Editor for .NET does not use any parts of the text processing stack from the OS ensures that document processing, measurement, layout, and visualization remain identical when the OS or .NET Core environment changes. Also, by solely utilizing managed memory and managed code, the control ensures that it is impervious to attacks using malicious fonts (for example, with malformed hinting code) or text sequences.

One DOM (to Rule Them All)

The Document Object Model (DOM) used by NOV Rich Text Editor for .NET is also managed and cross-platform. In fact, the entire text processing stack and text processor are implemented as managed .NET Standard/.NET Core assemblies, without dependencies on operating system services and with only minimal dependencies on .NET Core itself.

This ensures that services built around the DOM, such as history, serialization, and styling, are also platform-independent. The following list summarizes the tasks that can be performed using NOV Rich Text Editor for .NET on any operating system supported by the control:
  • Reports: Reports with embedded charts, barcodes, and drawings.
  • Business documents: Invoices, quotations, letters, and automatically generated contracts.
  • Document conversion: Load an existing document and export it to another format, such as DOCX to PDF, RTF to PDF, or RTF to DOCX.
  • Document editing: Let users review and revise generated content before saving or exporting it.

Document Formats

Building on the DOM, the control implements a set of managed, cross-platform readers and writers for commonly used document formats. These implementations are tightly integrated with NOV's managed Unicode and OpenType/TrueType implementations. For example, PDF font embedding uses information about the glyphs used in the document to embed only the required subset of font data, helping produce compact, self-contained PDF files. The following table lists the formats supported by NOV Rich Text Editor for .NET:

FormatImportExportPurpose
DOCXYesYesExchange formatted documents with Microsoft Word and other Office Open XML-compatible applications.
RTFYesYesExchange formatted text with applications supporting Rich Text Format.
HTMLYesYesLoad and save web content using supported HTML elements and CSS formatting.
EPUBYesYesGenerate electronic publications with HTML and CSS-based content.
TXTYesYesExchange plain text using supported Unicode and legacy encodings. Visual formatting is not retained.
PDFNoYesExport the laid-out document for distribution and printing.
NTX / NTBYesYesPreserve the entire NOV document state for storage and transfer. Both XML and binary versions are supported.


Conclusion

NOV Rich Text Editor provides a unique and consistent approach to cross-platform text processing in .NET. By implementing the entire text processing stack internally in completely managed code, the control ensures identical text processing and visualization across a wide range of platforms, including Windows, macOS, and Linux. It allows development teams to keep text processing separate from platform interaction and reuse their work in desktop applications, browser applications, and automated server-side services.

About Nevron Software

Founded in 1998, Nevron Software is a component vendor specializing in the development of premium presentation layer solutions for .NET-based technologies. Today, Nevron has established itself as a trusted partner worldwide for .NET LOB applications, SharePoint portals, and reporting solutions. Nevron technology is used by many Fortune 500 companies, large financial institutions, global IT consultancies, academic institutions, governments, and non-profits.
For more information, visit: www.nevron.com.
  
It may be of interest to you know that I selected Nevron charting component mostly based the examples demonstrating programmability, events, etc. We really didn't have the possibility to review in depth several graphing components (there are a lot of them) and all of them show pretty pictures, but from picture demos only you cannot tell about functionality and programmability inside the component engine.

Since Nevron examples demonstrated clicking, events, etc. I got the feeling that it offers a lot of freedom to the developer and you wouldn't be stuck with it at some later point in your development.
  

Kari Hirvi
TietoSaab Systems