ISO 32000-1 Document management — Portable document format — Part 1: PDF 1.7 - page 8

 

  Главная      Manuals     ISO 32000-1 Document management — Portable document format — Part 1: PDF 1.7

 

Search            copyright infringement  

 

 

 

 

 

 

 

 

 

 

 

Content      ..     6      7      8      9     ..

 

 

 

ISO 32000-1 Document management — Portable document format — Part 1: PDF 1.7 - page 8

 

 

9.7.5.2
Predefined CMaps
Several of the CMaps define mappings from Unicode encodings to character collections. Unicode values
appearing in a text string shall be represented in big-endian order (high-order byte first). CMap names
containing “UCS2” use UCS-2 encoding; names containing “UTF16” use UTF-16BE (big-endian) encoding.
NOTE 1
Table 118 lists the names of the predefined CMaps. These CMaps map character codes to CIDs in a single
descendant CIDFont. CMaps whose names end in H specify horizontal writing mode; those ending in V specify
vertical writing mode.
Table 118 - Predefined CJK CMap names
Name
Description
Chinese (Simplified)
GB-EUC-H
Microsoft Code Page 936 (lfCharSet 0x86), GB 2312-80 character set, EUC-CN
encoding
GB-EUC-V
Vertical version of GB-EUC-H
GBpc-EUC-H
Mac OS, GB 2312-80 character set, EUC-CN encoding, Script Manager code 19
GBpc-EUC-V
Vertical version of GBpc-EUC-H
GBK-EUC-H
Microsoft Code Page 936 (lfCharSet 0x86), GBK character set, GBK encoding
GBK-EUC-V
Vertical version of GBK-EUC-H
GBKp-EUC-H
Same as GBK-EUC-H but replaces half-width Latin characters with proportional
forms and maps character code 0x24 to a dollar sign ($) instead of a yuan symbol
(¥)
GBKp-EUC-V
Vertical version of GBKp-EUC-H
GBK2K-H
GB 18030-2000 character set, mixed 1-, 2-, and 4-byte encoding
GBK2K-V
Vertical version of GBK2K-H
UniGB-UCS2-H
Unicode (UCS-2) encoding for the Adobe-GB1 character collection
UniGB-UCS2-V
Vertical version of UniGB-UCS2-H
UniGB-UTF16-H
Unicode (UTF-16BE) encoding for the Adobe-GB1 character collection; contains
mappings for all characters in the GB18030-2000 character set
UniGB-UTF16-V
Vertical version of UniGB-UTF16-H
Chinese (Traditional)
B5pc-H
Mac OS, Big Five character set, Big Five encoding, Script Manager code 2
B5pc-V
Vertical version of B5pc-H
HKscs-B5-H
Hong Kong SCS, an extension to the Big Five character set and encoding
HKscs-B5-V
Vertical version of HKscs-B5-H
ETen-B5-H
Microsoft Code Page 950 (lfCharSet 0x88), Big Five character set with ETen
extensions
ETen-B5-V
Vertical version of ETen-B5-H
ETenms-B5-H
Same as ETen-B5-H but replaces half-width Latin characters with proportional
forms
273
Table 118 - Predefined CJK CMap names (continued)
Name
Description
ETenms-B5-V
Vertical version of ETenms-B5-H
CNS-EUC-H
CNS 11643-1992 character set, EUC-TW encoding
CNS-EUC-V
Vertical version of CNS-EUC-H
UniCNS-UCS2-H
Unicode (UCS-2) encoding for the Adobe-CNS1 character collection
UniCNS-UCS2-V
Vertical version of UniCNS-UCS2-H
UniCNS-UTF16-H
Unicode
(UTF-16BE) encoding for the Adobe-CNS1 character collection;
contains mappings for all the characters in the HKSCS-2001 character set and
contains both 2- and 4-byte character codes
UniCNS-UTF16-V
Vertical version of UniCNS-UTF16-H
Japanese
83pv-RKSJ-H
Mac OS, JIS X 0208 character set with KanjiTalk6 extensions, Shift-JIS encoding,
Script Manager code 1
90ms-RKSJ-H
Microsoft Code Page 932 (lfCharSet 0x80), JIS X 0208 character set with NEC
and IBM® extensions
90ms-RKSJ-V
Vertical version of 90ms-RKSJ-H
90msp-RKSJ-H
Same as 90ms-RKSJ-H but replaces half-width Latin characters with proportional
forms
90msp-RKSJ-V
Vertical version of 90msp-RKSJ-H
90pv-RKSJ-H
Mac OS, JIS X 0208 character set with KanjiTalk7 extensions, Shift-JIS encoding,
Script Manager code 1
Add-RKSJ-H
JIS X 0208 character set with Fujitsu FMR extensions, Shift-JIS encoding
Add-RKSJ-V
Vertical version of Add-RKSJ-H
EUC-H
JIS X 0208 character set, EUC-JP encoding
EUC-V
Vertical version of EUC-H
Ext-RKSJ-H
JIS C 6226 (JIS78) character set with NEC extensions, Shift-JIS encoding
Ext-RKSJ-V
Vertical version of Ext-RKSJ-H
H
JIS X 0208 character set, ISO-2022-JP encoding
V
Vertical version of H
UniJIS-UCS2-H
Unicode (UCS-2) encoding for the Adobe-Japan1 character collection
UniJIS-UCS2-V
Vertical version of UniJIS-UCS2-H
UniJIS-UCS2-HW-H
Same as UniJIS-UCS2-H but replaces proportional Latin characters with half-
width forms
UniJIS-UCS2-HW-V
Vertical version of UniJIS-UCS2-HW-H
UniJIS-UTF16-H
Unicode
(UTF-16BE) encoding for the Adobe-Japan1 character collection;
contains mappings for all characters in the JIS X 0213:1000 character set
UniJIS-UTF16-V
Vertical version of UniJIS-UTF16-H
Korean
274
Table 118 - Predefined CJK CMap names (continued)
Name
Description
KSC-EUC-H
KS X 1001:1992 character set, EUC-KR encoding
KSC-EUC-V
Vertical version of KSC-EUC-H
KSCms-UHC-H
Microsoft Code Page 949 (lfCharSet 0x81), KS X 1001:1992 character set plus
8822 additional hangul, Unified Hangul Code (UHC) encoding
KSCms-UHC-V
Vertical version of KSCms−UHC-H
KSCms-UHC-HW-H
Same as KSCms-UHC-H but replaces proportional Latin characters with half-
width forms
KSCms-UHC-HW-V
Vertical version of KSCms-UHC-HW-H
KSCpc-EUC-H
Mac OS, KS X 1001:1992 character set with Mac OS KH extensions, Script
Manager Code 3
UniKS-UCS2-H
Unicode (UCS-2) encoding for the Adobe-Korea1 character collection
UniKS-UCS2-V
Vertical version of UniKS-UCS2-H
UniKS-UTF16-H
Unicode (UTF-16BE) encoding for the Adobe-Korea1 character collection
UniKS-UTF16-V
Vertical version of UniKS-UTF16-H
Generic
Identity-H
The horizontal identity mapping for 2-byte CIDs; may be used with CIDFonts
using any Registry, Ordering, and Supplement values. It maps 2-byte character
codes ranging from 0 to 65,535 to the same 2-byte CID value, interpreted high-
order byte first.
Identity-V
Vertical version of Identity-H. The mapping is the same as for Identity-H.
NOTE 2
The Identity-H and Identity-V CMaps may be used to refer to glyphs directly by their CIDs when showing a text
string.
When the current font is a Type 0 font whose Encoding entry is Identity-H or Identity-V, the string to be shown
shall contain pairs of bytes representing CIDs, high-order byte first. When the current font is a CIDFont, the
string to be shown shall contain pairs of bytes representing CIDs, high-order byte first. When the current font is
a Type 2 CIDFont in which the CIDToGIDMap entry is Identity and if the TrueType font is embedded in the PDF
file, the
2-byte CID values shall be identical glyph indices for the glyph descriptions in the TrueType font
program.
NOTE 3
Table 119 lists the character collections referenced by the predefined CMaps for the different versions of PDF.
A dash (—) indicates that the CMap is not predefined in that PDF version.
Table 119 - Character collections for predefined CMaps, by PDF version
CMAP
PDF 1.2
PDF 1.3
PDF 1.4
PDF 1.5
Chinese (Simplified)
GB-EUC-H/V
Adobe-GB1-0
Adobe-GB1-0
Adobe-GB1-0
Adobe-GB1-0
GBpc-EUC-H
Adobe-GB1-0
Adobe-GB1-0
Adobe-GB1-0
Adobe-GB1-0
GBpc-EUC-V
Adobe-GB1-0
Adobe-GB1-0
Adobe-GB1-0
GBK-EUC-H/V
Adobe-GB1-2
Adobe-GB1-2
Adobe-GB1-2
GBKp-EUC-H/V
Adobe-GB1-2
Adobe-GB1-2
275
Table 119 - Character collections for predefined CMaps, by PDF version (continued)
CMAP
PDF 1.2
PDF 1.3
PDF 1.4
PDF 1.5
GBK2K-H/V
Adobe-GB1-4
Adobe-GB1-4
UniGB-UCS2-H/V
Adobe-GB1-2
Adobe-GB1-4
Adobe-GB1-4
UniGB-UTF16-H/V
Adobe-GB1-4
Chinese (Traditional)
B5pc-H/V
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
HKscs-B5-H/V
Adobe-CNS1-3
Adobe-CNS1-3
ETen-B5-H/V
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
ETenms-B5-H/V
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
CNS-EUC-H/V
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
Adobe-CNS1-0
UniCNS-UCS2-H/V
Adobe-CNS1-0
Adobe-CNS1-3
Adobe-CNS1-3
UniCNS-UTF16-H/V
Adobe-CNS1-4
Japanese
83pv-RKSJ-H
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
90ms-RKSJ-H/V
Adobe-Japan1-2
Adobe-Japan1-2
Adobe-Japan1-2
Adobe-Japan1-2
90msp-RKSJ-H/V
Adobe-Japan1-2
Adobe-Japan1-2
Adobe-Japan1-2
90pv-RKSJ-H
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
Add-RKSJ-H/V
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
EUC-H/V
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
Ext-RKSJ-H/V
Adobe-Japan1-2
Adobe-Japan1-2
Adobe-Japan1-2
Adobe-Japan1-2
H/V
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
Adobe-Japan1-1
UniJIS-UCS2-H/V
Adobe-Japan1-2
Adobe-Japan1-4
Adobe-Japan1-4
UniJIS-UCS2-HW-H/V
Adobe-Japan1-2
Adobe-Japan1-4
Adobe-Japan1-4
UniJIS-UTF16-H/V
Adobe-Japan1-5
Korean
KSC-EUC-H/V
Adobe-Korea1-0
Adobe-Korea1-0
Adobe-Korea1-0
Adobe-Korea1-0
KSCms-UHC-H/V
Adobe-Korea1-1
Adobe-Korea1-1
Adobe-Korea1-1
Adobe-Korea1-1
KSCms-UHC-HW-H/V
Adobe-Korea1-1
Adobe-Korea1-1
Adobe-Korea1-1
KSCpc-EUC-H
Adobe-Korea1-0
Adobe-Korea1-0
Adobe-Korea1-0
Adobe-Korea1-0
UniKS-UCS2-H/V
Adobe-Korea1-1
Adobe-Korea1-1
Adobe-Korea1-1
UniKS-UTF16-H/V
Adobe-Korea1-2
Generic
Identity-H/V
Adobe-Identity-0
Adobe-Identity-0
Adobe-Identity-0
Adobe-Identity-0
276
A conforming reader shall support all of the character collections listed in Table 119. As noted in 9.7.3,
"CIDSystemInfo Dictionaries", a character collection is identified by registry, ordering, and supplement number,
and supplements are cumulative; that is, a higher-numbered supplement includes the CIDs contained in lower-
numbered supplements, as well as some additional CIDs. Consequently, text encoded according to the
predefined CMaps for a given PDF version shall be valid when interpreted by a conforming reader supporting
the same or a later PDF version. When interpreted by a conforming reader supporting an earlier PDF version,
such text causes an error if a CMap is encountered that is not predefined for that PDF version. If character
codes are encountered that were added in a higher-numbered supplement than the one corresponding to the
supported PDF version, no characters are displayed for those codes; see 9.7.6.3, "Handling Undefined
Characters".
The Identity-H and Identity-V CMaps shall not be used with a non-embedded font. Only standardized character
sets may be used.
NOTE 4
If a conforming writer producing a PDF file encounters text to be included that uses CIDs from a higher-
numbered supplement than the one corresponding to the PDF version being generated, the application should
embed the CMap for the higher-numbered supplement rather than refer to the predefined CMap.
The CMap programs that define the predefined CMaps are available through the ASN Web site.
9.7.5.3
Embedded CMap Files
For character encodings that are not predefined, the PDF file shall contain a stream that defines the CMap. In
addition to the standard entries for streams (listed in Table 5), the CMap stream dictionary contains the entries
listed in Table 120. The data in the stream defines the mapping from character codes to a font number and a
character selector. The data shall follow the syntax defined in Adobe Technical Note #5014, Adobe CMap and
CIDFont Files Specification (see bibliography).
Table 120 - Additional entries in a CMap stream dictionary
Key
Type
Value
Type
name
(Required) The type of PDF object that this dictionary describes; shall
be CMap for a CMap dictionary.
CMapName
name
(Required) The name of the CMap. It shall be the same as the value of
CMapName in the CMap file.
CIDSystemInfo
dictionary
(Required) A dictionary
(see
9.7.3, "CIDSystemInfo Dictionaries")
containing entries that define the character collection for the CIDFont
or CIDFonts associated with the CMap.
The value of this entry shall be the same as the value of
CIDSystemInfo in the CMap file. (However, it does not need to match
the values of CIDSystemInfo for the Identity-H or Identity-V CMaps.)
WMode
integer
(Optional) A code that specifies the writing mode for any CIDFont with
which this CMap is combined. The value shall be 0 for horizontal or 1
for vertical. Default value: 0.
The value of this entry shall be the same as the value of WMode in the
CMap file.
UseCMap
name or
(Optional) The name of a predefined CMap, or a stream containing a
stream
CMap. If this entry is present, the referencing CMap shall specify only
the character mappings that differ from the referenced CMap.
9.7.5.4
CMap Example and Operator Summary
Embedded CMap files shall conform to the format documented in Adobe Technical Note #5014, subject to
these additional constraints:
277
a) If the embedded CMap file contains a usecmap reference, the CMap indicated there shall also be identified
by the UseCMap entry in the CMap stream dictionary.
b) The usefont operator, if present, shall specify a font number of 0.
c) The beginbfchar and endbfchar shall not appear in a CMap that is used as the Encoding entry of a Type 0
font; however, they may appear in the definition of a ToUnicode CMap.
d) A notdef mapping, defined using beginnotdefchar, endnotdefchar, beginnotdefrange, and endnotdefrange
shall be used if the normal mapping produces a CID for which no glyph is present in the associated
CIDFont.
e) The beginrearrangedfont, endrearrangedfont, beginusematrix, and endusematrix operators shall not be
used.
EXAMPLE
This example shows a sample CMap for a Japanese Shift-JIS encoding. Character codes in this encoding
can be either 1 or 2 bytes in length. This CMap could be used with a CIDFont that uses the same CID
ordering as specified in the CIDSystemInfo entry. Note that several of the entries in the stream dictionary
are also replicated in the stream data.
22 0 obj
<<
/Type /CMap
/CMapName /90ms-RKSJ-H
/CIDSystemInfo <<
/Registry
( Adobe )
/Ordering ( Japan1 )
/Supplement 2
>>
/WMode 0
/Length 23 0 R
>>
stream
%!PS-Adobe-3 . 0 Resource-CMap
%%DocumentNeededResources : ProcSet ( CIDInit )
%%IncludeResource : ProcSet ( CIDInit )
%%BeginResource : CMap ( 90ms-RKSJ-H )
%%Title : ( 90ms-RKSJ-H Adobe Japan1 2 )
%%Version : 10 . 001
%%Copyright : Copyright 1990-2001 Adobe Systems Inc .
%%Copyright : All Rights Reserved .
%%EndComments
/CIDInit
/ProcSet findresource begin
12 dict begin
begincmap
/CIDSystemInfo
3 dict dup begin
/Registry
( Adobe ) def
/Ordering ( Japan1 ) def
/Supplement 2 def
end def
/CMapName /90ms-RKSJ-H def
/CMapVersion 10 . 001 def
/CMapType 1 def
/UIDOffset 950 def
/XUID [ 1 10 25343 ] def
/WMode 0 def
4 begincodespacerange
< 00 >
< 80 >
< 8140 >
< 9FFC >
< A0 >
< DF >
278
< E040 >
< FCFC >
endcodespacerange
1 beginnotdefrange
< 00 >
< 1F >
231
endnotdefrange
100 begincidrange
< 20 >
< 7D >
231
< 7E >
< 7E >
631
< 8140 >
< 817E >
633
< 8180 >
< 81AC >
696
< 81B8 >
< 81BF >
741
< 81C8 >
< 81CE >
749
… Additional ranges…
< FB40 >
< FB7E >
8518
< FB80 >
< FBFC >
8581
< FC40 >
< FC4B >
8706
endcidrange
endcmap
CMapName currentdict /CMap defineresource pop
end
end
%%EndResource
%%EOF
endstream
endobj
9.7.6
Type 0 Font Dictionaries
9.7.6.1
General
A Type 0 font dictionary contains the entries listed in Table 121.
Table 121 - Entries in a Type 0 font dictionary
Key
Type
Value
Type
name
(Required) The type of PDF object that this dictionary describes; shall
be Font for a font dictionary.
Subtype
name
(Required) The type of font; shall be Type0 for a Type 0 font.
BaseFont
name
(Required) The name of the font. If the descendant is a Type 0
CIDFont, this name should be the concatenation of the CIDFont’s
BaseFont name, a hyphen, and the CMap name given in the
Encoding entry (or the CMapName entry in the CMap). If the
descendant is a Type 2 CIDFont, this name should be the same as the
CIDFont’s BaseFont name.
NOTE
In principle, this is an arbitrary name, since there is no
font program associated directly with a Type
0 font
dictionary. The conventions described here ensure
maximum compatibility with existing readers.
Encoding
name or
(Required) The name of a predefined CMap, or a stream containing a
stream
CMap that maps character codes to font numbers and CIDs. If the
descendant is a Type 2 CIDFont whose associated TrueType font
program is not embedded in the PDF file, the Encoding entry shall be
a predefined CMap name (see 9.7.4.2, "Glyph Selection in CIDFonts").
279
Table 121 - Entries in a Type 0 font dictionary (continued)
Key
Type
Value
DescendantFonts
array
(Required) A one-element array specifying the CIDFont dictionary that
is the descendant of this Type 0 font.
ToUnicode
stream
(Optional) A stream containing a CMap file that maps character codes
to Unicode values (see 9.10, "Extraction of Text Content").
EXAMPLE
This code sample shows a Type 0 font.
14 0 obj
<<
/Type /Font
/Subtype /Type0
/BaseFont /HeiseiMin-W5-90ms-RKSJ-H
/Encoding /90ms-RKSJ-H
/DescendantFonts [ 15 0 R ]
>>
endobj
9.7.6.2
CMap Mapping
The Encoding entry of a Type 0 font dictionary specifies a CMap that specifies how text-showing operators
(such as Tj) shall interpret the bytes in the string to be shown when the current font is the Type 0 font. This sub-
clause describes how the characters in the string shall be decoded and mapped into character selectors, which
in PDF are always CIDs.
The codespace ranges in the CMap (delimited by begincodespacerange and endcodespacerange) specify
how many bytes are extracted from the string for each successive character code. A codespace range shall be
specified by a pair of codes of some particular length giving the lower and upper bounds of that range. A code
shall be considered to match the range if it is the same length as the bounding codes and the value of each of
its bytes lies between the corresponding bytes of the lower and upper bounds. The code length shall not be
greater than 4.
A sequence of one or more bytes shall be extracted from the string and matched against the codespace ranges
in the CMap. That is, the first byte shall be matched against 1-byte codespace ranges; if no match is found, a
second byte shall be extracted, and the 2-byte code shall be matched against 2-byte codespace ranges. This
process continues for successively longer codes until a match is found or all codespace ranges have been
tested. There will be at most one match because codespace ranges shall not overlap.
The code extracted from the string shall be looked up in the character code mappings for codes of that length.
(These are the mappings defined by beginbfchar, endbfchar, begincidchar, endcidchar, and corresponding
operators for ranges.) Failing that, it shall be looked up in the notdef mappings, as described in the next sub-
clause.
The results of the CMap mapping algorithm are a font number and a character selector. The font number shall
be used as an index into the Type 0 font’s DescendantFonts array to select a CIDFont. In PDF, the font
number shall be 0 and the character selector shall be a CID; this is the only case described here. The CID shall
then be used to select a glyph in the CIDFont. If the CIDFont contains no glyph for that CID, the notdef
mappings shall be consulted, as described in 9.7.6.3, "Handling Undefined Characters".
9.7.6.3
Handling Undefined Characters
A CMap mapping operation can fail to select a glyph for a variety of reasons. This sub-clause describes those
reasons and what happens when they occur.
If a code maps to a CID for which no such glyph exists in the descendant CIDFont, the notdef mappings in the
CMap shall be consulted to obtain a substitute character selector. These mappings are delimited by the
280
operators beginnotdefchar, endnotdefchar, beginnotdefrange, and endnotdefrange within an embedded
CMap file. They shall always map to a CID. If a matching notdef mapping is found, the CID selects a glyph in
the associated descendant, which shall be a CIDFont. If no glyph exists for that CID, the glyph for CID 0 (which
shall be present) shall be substituted.
NOTE 5
The notdef mappings are similar to the . notdef character mechanism in simple fonts.
If the CMap does not contain either a character mapping or a notdef mapping for the code, descendant 0 shall
be selected and the glyph for CID 0 shall be substituted from the associated CIDFont.
If the code is invalid—that is, the bytes extracted from the string to be shown do not match any codespace
range in the CMap—a substitute glyph is chosen as just described. The character mapping algorithm shall be
reset to its original position in the string, and a modified mapping algorithm chooses the best partially matching
codespace range:
a) If the first byte extracted from the string to be shown does not match the first byte of any codespace range,
the range having the shortest codes shall be chosen.
b) Otherwise (that is, if there is a partial match), for each additional byte extracted, the code accumulated so
far shall be matched against the beginnings of all longer codespace ranges until the longest such partial
match has been found. If multiple codespace ranges have partial matches of the same length, the one
having the shortest codes shall be chosen.
The length of the codes in the chosen codespace range determines the total number of bytes to consume from
the string for the current mapping operation.
9.8
Font Descriptors
9.8.1
General
A font descriptor specifies metrics and other attributes of a simple font or a CIDFont as a whole, as distinct from
the metrics of individual glyphs. These font metrics provide information that enables a conforming reader to
synthesize a substitute font or select a similar font when the font program is unavailable. The font descriptor
may also be used to embed the font program in the PDF file.
Font descriptors shall not be used with Type 0 fonts. Beginning with PDF 1.5, font descriptors may be used with
Type 3 fonts.
A font descriptor is a dictionary whose entries specify various font attributes. The entries common to all font
descriptors—for both simple fonts and CIDFonts—are listed in Table 122. Additional entries in the font
descriptor for a CIDFont are described in 9.8.3, "Font Descriptors for CIDFonts". All integer values shall be
units in glyph space. The conversion from glyph space to text space is described in 9.2.4, "Glyph Positioning
and Metrics".
Table 122 - Entries common to all font descriptors
Key
Type
Value
Type
name
(Required) The type of PDF object that this dictionary describes; shall
be FontDescriptor for a font descriptor.
FontName
name
(Required) The PostScript name of the font. This name shall be the
same as the value of BaseFont in the font or CIDFont dictionary that
refers to this font descriptor.
FontFamily
byte string
(Optional; PDF 1.5; should be used for Type 3 fonts in Tagged PDF
documents) A byte string specifying the preferred font family name.
EXAMPLE 1
For the font Times Bold Italic, the FontFamily is
Times.
281
Table 122 - Entries common to all font descriptors (continued)
Key
Type
Value
FontStretch
name
(Optional; PDF 1.5; should be used for Type 3 fonts in Tagged PDF
documents) The font stretch value. It shall be one of these names
(ordered
from
narrowest
to
widest):
UltraCondensed,
ExtraCondensed,
Condensed,
SemiCondensed,
Normal,
SemiExpanded, Expanded, ExtraExpanded or UltraExpanded.
The specific interpretation of these values varies from font to font.
EXAMPLE 2
Condensed in one font may appear most similar to
Normal in another.
FontWeight
number
(Optional; PDF 1.5; should be used for Type 3 fonts in Tagged PDF
documents) The weight (thickness) component of the fully-qualified
font name or font specifier. The possible values shall be 100, 200, 300,
400, 500, 600, 700, 800, or 900, where each number indicates a
weight that is at least as dark as its predecessor. A value of 400 shall
indicate a normal weight; 700 shall indicate bold.
The specific interpretation of these values varies from font to font.
EXAMPLE 3
300 in one font may appear most similar to 500 in
another.
Flags
integer
(Required) A collection of flags defining various characteristics of the
font (see 9.8.2, "Font Descriptor Flags").
FontBBox
rectangle
(Required, except for Type
3 fonts) A rectangle
(see
7.9.5,
"Rectangles"), expressed in the glyph coordinate system, that shall
specify the font bounding box. This should be the smallest rectangle
enclosing the shape that would result if all of the glyphs of the font
were placed with their origins coincident and then filled.
ItalicAngle
number
(Required) The angle, expressed in degrees counterclockwise from
the vertical, of the dominant vertical strokes of the font.
EXAMPLE 4
The 9-o’clock position is 90 degrees, and the 3-
o’clock position is -90 degrees.
The value shall be negative for fonts that slope to the right, as almost
all italic fonts do.
Ascent
number
(Required, except for Type 3 fonts) The maximum height above the
baseline reached by glyphs in this font. The height of glyphs for
accented characters shall be excluded.
Descent
number
(Required, except for Type 3 fonts) The maximum depth below the
baseline reached by glyphs in this font. The value shall be a negative
number.
Leading
number
(Optional) The spacing between baselines of consecutive lines of text.
Default value: 0.
CapHeight
number
(Required for fonts that have Latin characters, except for Type 3 fonts)
The vertical coordinate of the top of flat capital letters, measured from
the baseline.
XHeight
number
(Optional) The font’s x height: the vertical coordinate of the top of flat
nonascending lowercase letters (like the letter x), measured from the
baseline, in fonts that have Latin characters. Default value: 0.
StemV
number
(Required, except for Type
3 fonts) The thickness, measured
horizontally, of the dominant vertical stems of glyphs in the font.
StemH
number
(Optional) The thickness, measured vertically, of the dominant
horizontal stems of glyphs in the font. Default value: 0.
282
Table 122 - Entries common to all font descriptors (continued)
Key
Type
Value
AvgWidth
number
(Optional) The average width of glyphs in the font. Default value: 0.
MaxWidth
number
(Optional) The maximum width of glyphs in the font. Default value: 0.
MissingWidth
number
(Optional) The width to use for character codes whose widths are not
specified in a font dictionary’s Widths array. This shall have a
predictable effect only if all such codes map to glyphs whose actual
widths are the same as the value of the MissingWidth entry. Default
value: 0.
FontFile
stream
(Optional) A stream containing a Type
1 font program (see 9.9,
"Embedded Font Programs").
FontFile2
stream
(Optional; PDF 1.1) A stream containing a TrueType font program (see
9.9, "Embedded Font Programs").
FontFile3
stream
(Optional; PDF 1.2) A stream containing a font program whose format
is specified by the Subtype entry in the stream dictionary
(see
Table 126).
CharSet
ASCII string
(Optional; meaningful only in Type 1 fonts; PDF 1.1) A string listing the
or
byte
character names defined in a font subset. The names in this string
string
shall be in PDF syntax—that is, each name preceded by a slash (/).
The names may appear in any order. The name . notdef shall be
omitted; it shall exist in the font subset. If this entry is absent, the only
indication of a font subset shall be the subset tag in the FontName
entry (see 9.6.4, "Font Subsets").
At most, only one of the FontFile, FontFile2, and FontFile3 entries shall be present.
9.8.2
Font Descriptor Flags
The value of the Flags entry in a font descriptor shall be an unsigned 32-bit integer containing flags specifying
various characteristics of the font. Bit positions within the flag word are numbered from 1 (low-order) to 32
(high-order). Table 123 shows the meanings of the flags; all undefined flag bits are reserved and shall be set to
0 by conforming writers. Figure 48 shows examples of fonts with these characteristics.
Table 123 - Font flags
Bit position
Name
Meaning
1
FixedPitch
All glyphs have the same width
(as opposed to proportional or
variable-pitch fonts, which have different widths).
2
Serif
Glyphs have serifs, which are short strokes drawn at an angle on the
top and bottom of glyph stems. (Sans serif fonts do not have serifs.)
3
Symbolic
Font contains glyphs outside the Adobe standard Latin character set.
This flag and the Nonsymbolic flag shall not both be set or both be
clear.
4
Script
Glyphs resemble cursive handwriting.
6
Nonsymbolic
Font uses the Adobe standard Latin character set or a subset of it.
7
Italic
Glyphs have dominant vertical strokes that are slanted.
17
AllCap
Font contains no lowercase letters; typically used for display purposes,
such as for titles or headlines.
283
Table 123 - Font flags (continued)
Bit position
Name
Meaning
18
SmallCap
Font contains both uppercase and lowercase letters. The uppercase
letters are similar to those in the regular version of the same typeface
family. The glyphs for the lowercase letters have the same shapes as
the corresponding uppercase letters, but they are sized and their
proportions adjusted so that they have the same size and stroke
weight as lowercase glyphs in the same typeface family.
19
ForceBold
See description after Note 1 in this sub-clause.
The Nonsymbolic flag (bit 6 in the Flags entry) indicates that the font’s character set is the Adobe standard
Latin character set (or a subset of it) and that it uses the standard names for those glyphs. This character set is
shown in D.2, "Latin Character Set and Encodings". If the font contains any glyphs outside this set, the
Symbolic flag shall be set and the Nonsymbolic flag shall be clear. In other words, any font whose character set
is not a subset of the Adobe standard character set shall be considered to be symbolic. This influences the
font’s implicit base encoding and may affect a conforming reader’s font substitution strategies.
Fixed-pitch font
The quick brown fox jumped.
Serif font
The quick brown fox jumped.
Sans serif font
The quick brown fox jumped.
Symbolic font
✴❈❅ ❑◆❉❃❋ ❂❒❏◗■ ❆❏❘ ❊◆❍❐❅❄✎
Script font
The quick brown fox jumped.
Italic font
The quick brown fox jumped.
All-cap font
The quick brown fox jumped
Small-cap font
The quick brown fox jumped.
Figure 48 - Characteristics represented in the Flags entry of a font descriptor
NOTE 1
This classification of nonsymbolic and symbolic fonts is peculiar to PDF. A font may contain additional
characters that are used in Latin writing systems but are outside the Adobe standard Latin character set; PDF
considers such a font to be symbolic. The use of two flags to represent a single binary choice is a historical
accident.
The ForceBold flag (bit 19) shall determine whether bold glyphs shall be painted with extra pixels even at very
small text sizes by a conforming reader. If the ForceBold flag is set, features of bold glyphs may be thickened at
small text sizes.
NOTE 2
Typically, when glyphs are painted at small sizes on very low-resolution devices such as display screens,
features of bold glyphs may appear only 1 pixel wide. Because this is the minimum feature width on a pixel-
based device, ordinary (nonbold) glyphs also appear with 1-pixel-wide features and therefore cannot be
distinguished from bold glyphs.
284
EXAMPLE
This code sample illustrates a font descriptor whose Flags entry has the Serif, Nonsymbolic, and
ForceBold flags (bits 2, 6, and 19) set.
7 0 obj
<<
/Type /FontDescriptor
/FontName /AGaramond-Semibold
/Flags 262178
% Bits 2, 6, and 19
/FontBBox [ −177 −269 1123 866 ]
/MissingWidth 255
/StemV 105
/StemH 45
/CapHeight 660
/XHeight 394
/Ascent 720
/Descent −270
/Leading 83
/MaxWidth 1212
/AvgWidth 478
/ItalicAngle
0
>>
endobj
9.8.3
Font Descriptors for CIDFonts
9.8.3.1
General
In addition to the entries in Table 122, the FontDescriptor dictionaries of CIDFonts may contain the entries
listed in Table 124.
Table 124 - Additional font descriptor entries for CIDFonts
Key
Type
Value
Style
dictionary
(Optional) A dictionary containing entries that describe the style of the glyphs in
the font (see 9.8.3.2, "Style").
Lang
name
(Optional; PDF 1.5) A name specifying the language of the font, which may be
used for encodings where the language is not implied by the encoding itself.
The value shall be one of the codes defined by Internet RFC 3066, Tags for the
Identification of Languages or (PDF 1.0) 2-character language codes defined
by ISO 639 (see the Bibliography). If this entry is absent, the language shall be
considered to be unknown.
FD
dictionary
(Optional) A dictionary whose keys identify a class of glyphs in a CIDFont.
Each value shall be a dictionary containing entries that shall override the
corresponding values in the main font descriptor dictionary for that class of
glyphs (see 9.8.3.3, "FD").
CIDSet
stream
(Optional) A stream identifying which CIDs are present in the CIDFont file. If
this entry is present, the CIDFont shall contain only a subset of the glyphs in
the character collection defined by the CIDSystemInfo dictionary. If it is
absent, the only indication of a CIDFont subset shall be the subset tag in the
FontName entry (see 9.6.4, "Font Subsets").
The stream’s data shall be organized as a table of bits indexed by CID. The bits
shall be stored in bytes with the high-order bit first. Each bit shall correspond to
a CID. The most significant bit of the first byte shall correspond to CID 0, the
next bit to CID 1, and so on.
9.8.3.2
Style
The Style dictionary contains entries that define style attributes and values for the CIDFont. Only the Panose
entry is defined. The value of Panose shall be a 12-byte string consisting of these elements:
285
The font family class and subclass ID bytes, given in the sFamilyClass field of the “OS/2” table in a
TrueType font. This field is documented in Microsoft’s TrueType 1.0 Font Files Technical Specification.
Ten bytes for the PANOSE classification number for the font. The PANOSE classification system is
documented in Hewlett-Packard Company’s PANOSE Classification Metrics Guide.
See the Bibliography for more information about these documents.
EXAMPLE
This is an example of a Style entry in the font descriptor:
/Style
<< /Panose < 01 05 02 02 03 00 00 00 00 00 00 00 > >>
9.8.3.3
FD
A CIDFont may be made up of different classes of glyphs, each class requiring different sets of the font-wide
attributes that appear in font descriptors.
EXAMPLE 1
Latin glyphs, for example, may require different attributes than kanji glyphs.
The font descriptor shall define a set of default attributes that apply to all glyphs in the CIDFont. The FD entry in
the font descriptor shall contain exceptions to these defaults.
The key for each entry in an FD dictionary shall be the name of a class of glyphs—that is, a particular subset of
the CIDFont’s character collection. The entry’s value shall be a font descriptor whose contents shall override
the font-wide attributes for that class only. This font descriptor shall contain entries for metric information only; it
shall not include FontFile, FontFile2, FontFile3, or any of the entries listed in Table 122.
The FD dictionary should contain at least the metrics for the proportional Latin glyphs. With the information for
these glyphs, a more accurate substitution font can be created.
The names of the glyph classes depend on the character collection, as identified by the Registry, Ordering,
and Supplement entries in the CIDSystemInfo dictionary. Table 125 lists the valid keys for the Adobe-GB1,
Adobe-CNS1, Adobe-Japan1, Adobe-Japan2, and Adobe-Korea1 character collections.
286
Table 125 - Glyph classes in CJK fonts
Character Collection
Class
Glyphs in Class
Adobe-GB1
Alphabetic
Full-width Latin, Greek, and Cyrillic glyphs
Dingbats
Special symbols
Generic
Typeface-independent glyphs, such as line-drawing
Hanzi
Full-width hanzi (Chinese) glyphs
HRoman
Half-width Latin glyphs
HRomanRot
Same as HRoman but rotated for use in vertical writing
Kana
Japanese kana (katakana and hiragana) glyphs
Proportional
Proportional Latin glyphs
ProportionalRot
Same as Proportional but rotated for use in vertical
writing
Adobe-CNS1
Alphabetic
Full-width Latin, Greek, and Cyrillic glyphs
Dingbats
Special symbols
Generic
Typeface-independent glyphs, such as line-drawing
Hanzi
Full-width hanzi (Chinese) glyphs
HRoman
Half-width Latin glyphs
HRomanRot
Same as HRoman but rotated for use in vertical writing
Kana
Japanese kana (katakana and hiragana) glyphs
Proportional
Proportional Latin glyphs
ProportionalRot
Same as Proportional but rotated for use in vertical
writing
Adobe-Japan1
Alphabetic
Full-width Latin, Greek, and Cyrillic glyphs
AlphaNum
Numeric glyphs
Dingbats
Special symbols
DingbatsRot
Same as Dingbats but rotated for use in vertical writing
Generic
Typeface-independent glyphs, such as line-drawing
GenericRot
Same as Generic but rotated for use in vertical writing
HKana
Half-width kana (katakana and hiragana) glyphs
HKanaRot
Same as HKana but rotated for use in vertical writing
HRoman
Half-width Latin glyphs
HRomanRot
Same as HRoman but rotated for use in vertical writing
Kana
Full-width kana (katakana and hiragana) glyphs
Kanji
Full-width kanji (Chinese) glyphs
Proportional
Proportional Latin glyphs
ProportionalRot
Same as Proportional but rotated for use in vertical
writing
Glyphs used for setting ruby (small glyphs that serve to
Ruby
annotate other glyphs with meanings or readings)
Adobe-Japan2
Alphabetic
Full-width Latin, Greek, and Cyrillic glyphs
Dingbats
Special symbols
HojoKanji
Full-width kanji glyphs
287
Table 125 - Glyph classes in CJK fonts (continued)
Character Collection
Class
Glyphs in Class
Adobe-Korea1
Alphabetic
Full-width Latin, Greek, and Cyrillic glyphs
Dingbats
Special symbols
Generic
Typeface-independent glyphs, such as line-drawing
Hangul
Hangul and jamo glyphs
Hanja
Full-width hanja (Chinese) glyphs
HRoman
Half-width Latin glyphs
HRomanRot
Same as HRoman but rotated for use in vertical writing
Kana
Japanese kana (katakana and hiragana) glyphs
Proportional
Proportional Latin glyphs
ProportionalRot
Same as Proportional but
rotated for
use
in vertical
writing
EXAMPLE 2
This example illustrates an FD dictionary containing two entries.
/FD
<<
/Proportional
25 0 R
/HKana 26 0 R
>>
25 0 obj
<<
/Type /FontDescriptor
/FontName /HeiseiMin-W3-Proportional
/Flags 2
/AvgWidth 478
/MaxWidth 1212
/MissingWidth 250
/StemV 105
/StemH 45
/CapHeight 660
/XHeight 394
/Ascent 720
/Descent −270
/Leading 83
>>
endobj
26 0 obj
<< /Type /FontDescriptor
/FontName /HeiseiMin-W3-HKana
/Flags 3
/AvgWidth 500
/MaxWidth 500
/MissingWidth 500
/StemV 50
/StemH 75
/Ascent 720
/Descent 0
/Leading 83
>>
endobj
9.9
Embedded Font Programs
A font program may be embedded in a PDF file as data contained in a PDF stream object.
NOTE 1
Such a stream object is also called a font file by analogy with font programs that are available from sources
external to the conforming writer.
288
Font programs are subject to copyright, and the copyright owner may impose conditions under which a font
program may be used. These permissions are recorded either in the font program or as part of a separate
license. One of the conditions may be that the font program cannot be embedded, in which case it should not
be incorporated into a PDF file. A font program may allow embedding for the sole purpose of viewing and
printing the document but not for creating new or modified text that uses the font (in either the same document
or other documents). The latter operation would require the user performing the operation to have a licensed
copy of the font program, not a copy extracted from the PDF file. In the absence of explicit information to the
contrary, embedded font programs shall be used only to view and print the document and not for any other
purposes.
Table 126 summarizes the ways in which font programs shall be embedded in a PDF file, depending on the
representation of the font program. The key shall be the name used in the font descriptor to refer to the font file
stream; the subtype shall be the value of the Subtype key, if present, in the font file stream dictionary. Further
details of specific font program representations are given below.
Table 126 - Embedded font organization for various font types
Key
Subtype
Description
FontFile
Type 1 font program, in the original (noncompact) format described
in Adobe Type 1 Font Format. This entry may appear in the font
descriptor for a Type1 or MMType1 font dictionary.
FontFile2
(PDF 1.1) TrueType font program, as described in the TrueType
Reference Manual. This entry may appear in the font descriptor for
a TrueType font dictionary or
(PDF 1.3) for a CIDFontType2
CIDFont dictionary.
FontFile3
Type1C
(PDF 1.2) Type
1-equivalent font program represented in the
Compact Font Format (CFF), as described in Adobe Technical Note
#5176, The Compact Font Format Specification. This entry may
appear in the font descriptor for a Type1 or MMType1 font
dictionary.
FontFile3
CIDFontType0C
(PDF 1.3) Type 0 CIDFont program represented in the Compact
Font Format (CFF), as described in Adobe Technical Note #5176,
The Compact Font Format Specification. This entry may appear in
the font descriptor for a CIDFontType0 CIDFont dictionary.
289
Table 126 - Embedded font organization for various font types (continued)
Key
Subtype
Description
FontFile3
OpenType
(PDF 1.6) OpenType® font program, as described in the OpenType
Specification v.1.4
(see the Bibliography). OpenType is an
extension of TrueType that allows inclusion of font programs that
use the Compact Font Format (CFF).
A FontFile3 entry with an OpenType subtype may appear in the
font descriptor for these types of font dictionaries:
• A TrueType font dictionary or a CIDFontType2 CIDFont
dictionary, if the embedded font program contains a “glyf” table.
In addition to the “glyf” table, the font program must include
these tables: “head”, “hhea”, “hmtx”, “loca”, and “maxp”. The
“cvt ” (notice the trailing SPACE), “fpgm”, and “prep” tables must
also be included if they are required by the font instructions.
• A CIDFontType0 CIDFont dictionary, if the embedded font
program contains a “CFF ” table (notice the trailing SPACE) with
a Top DICT that uses CIDFont operators (this is equivalent to
subtype CIDFontType0C). In addition to the “CFF ” table, the
font program must include the “cmap” table.
• A Type1 font dictionary or CIDFontType0 CIDFont dictionary, if
the embedded font program contains a “CFF ” table without
CIDFont operators. In addition to the “CFF ” table, the font
program must include the “cmap” table.
The OpenType Specification describes a set of required tables;
however, not all tables are required in the font file, as described for
each type of font dictionary that can include this entry.
NOTE
The absence of some optional tables (such as those
used for advanced line layout) may prevent editing of
text containing the font.
The stream dictionary for a font file shall contain the normal entries for a stream, such as Length and Filter
(listed in Table 5), plus the additional entries listed in Table 127.
Table 127 - Additional entries in an embedded font stream dictionary
Key
Type
Value
Length1
integer
(Required for Type 1 and TrueType fonts) The length in bytes of the clear-text
portion of the Type 1 font program, or the entire TrueType font program, after it has
been decoded using the filters specified by the stream’s Filter entry, if any.
Length2
integer
(Required for Type 1 fonts) The length in bytes of the encrypted portion of the Type
1 font program after it has been decoded using the filters specified by the stream’s
Filter entry.
Length3
integer
(Required for Type 1 fonts) The length in bytes of the fixed-content portion of the
Type 1 font program after it has been decoded using the filters specified by the
stream’s Filter entry. If Length3 is
0, it indicates that the
512 zeros and
cleartomark have not been included in the FontFile font program and shall be
added by the conforming reader.
Subtype
name
(Required if referenced from FontFile3; PDF 1.2) A name specifying the format of
the embedded font program. The name shall be Type1C for Type 1 compact fonts,
CIDFontType0C for Type 0 compact CIDFonts, or OpenType for OpenType fonts.
Metadata
stream
(Optional; PDF 1.4) A metadata stream containing metadata for the embedded font
program (see 14.3.2, "Metadata Streams").
NOTE 2
A standard Type 1 font program, as described in the Adobe Type 1 Font Format specification, consists of three
parts: a clear-text portion (written using PostScript syntax), an encrypted portion, and a fixed-content portion.
The fixed-content portion contains 512 ASCII zeros followed by a cleartomark operator, and perhaps followed
290
by additional data. Although the encrypted portion of a standard Type 1 font may be in binary or ASCII
hexadecimal format, PDF supports only the binary format. However, the entire font program may be encoded
using any filters.
EXAMPLE
This code shows the structure of an embedded standard Type 1 font.
12 0 obj
<<
/Filter
/ASCII85Decode
/Length 41116
/Length1 2526
/Length2 32393
/Length3 570
>>
stream
,p>`rDKJj'E+LaU0eP.@+AH9dBOu$hFD55nC
Omitted data
JJQ&Nt')<=^p&mGf(%:%h1%9c//K(/*o=.C>UXkbVGTrr~>
endstream
endobj
As noted in Table 126, a Type 1-equivalent font program or a Type 0 CIDFont program may be represented in
the Compact Font Format (CFF). The Length1, Length2, and Length3 entries are not needed in that case and
shall not be present. Although CFF enables multiple font or CIDFont programs to be bundled together in a
single file, an embedded CFF font file in PDF shall consist of exactly one font or CIDFont (as appropriate for the
associated font dictionary).
According to the Adobe Type 1 Font Format specification, a Type 1 font program may contain a PaintType
entry specifying whether the glyphs’ outlines are to be filled or stroked. For fonts embedded in a PDF file, this
entry shall be ignored; the decision whether to fill or stroke glyph outlines is entirely determined by the PDF text
rendering mode parameter (see 9.3.6, "Text Rendering Mode"). This shall also applies to Type 1 compact fonts
and Type 0 compact CIDFonts.
A TrueType font program may be used as part of either a font or a CIDFont. Although the basic font file format
is the same in both cases, there are different requirements for what information shall be present in the font
program. These TrueType tables shall always be present if present in the original TrueType font program:
“head”, “hhea”, “loca”, “maxp”, “cvt”, “prep”, “glyf”, “hmtx”, and “fpgm”. If used with a simple font dictionary, the
font program shall additionally contain a cmap table defining one or more encodings, as discussed in 9.6.6.4,
"Encodings for TrueType Fonts". If used with a CIDFont dictionary, the cmap table is not needed and shall not
be present, since the mapping from character codes to glyph descriptions is provided separately.
The “vhea” and “vmtx” tables that specify vertical metrics shall never be used by a conforming reader. The only
way to specify vertical metrics in PDF shall be by means of the DW2 and W2 entries in a CIDFont dictionary.
NOTE 3
Beginning with PDF 1.6, font programs may be embedded using the OpenType format, which is an extension
of the TrueType format that allows inclusion of font programs using the Compact Font Format (CFF). It also
allows inclusion of data to describe glyph substitutions, kerning, and baseline adjustments. In addition to
rendering glyphs, conforming readers may use the data in OpenType fonts to do advanced line layout,
automatically substitute ligatures, provide selections of alternate glyphs to users, and handle complex writing
scripts.
The process of finding glyph descriptions in OpenType fonts by a conforming reader shall be the following:
For Type 1 fonts using “CFF” tables, the process shall be as described in 9.6.6.2, "Encodings for Type 1
Fonts".
For TrueType fonts using “glyf” tables, the process shall be as described in 9.6.6.4, "Encodings for
TrueType Fonts". Since this process sometimes produces ambiguous results, conforming writers, instead
of using a simple font, shall use a Type 0 font with an Identity-H encoding and use the glyph indices as
character codes, as described following Table 118.
291
For CIDFontType0 fonts using “CFF” tables, the process shall be as described in the discussion of
embedded Type 0 CIDFonts in 9.7.4.2, "Glyph Selection in CIDFonts".
For CIDFontType2 fonts using “glyf” tables, the process shall be as described in the discussion of
embedded Type 2 CIDFonts in 9.7.4.2, "Glyph Selection in CIDFonts".
As discussed in 9.6.4, "Font Subsets", an embedded font program may contain only the subset of glyphs that
are used in the PDF document. This may be indicated by the presence of a CharSet or CIDSet entry in the font
descriptor that refers to the font file.
9.10
Extraction of Text Content
9.10.1
General
The preceding sub-clauses describe all the facilities for showing text and causing glyphs to be painted on the
page. In addition to displaying text, conforming readers sometimes need to determine the information content
of text—that is, its meaning according to some standard character identification as opposed to its rendered
appearance. This need arises during operations such as searching, indexing, and exporting of text to other file
formats.
The Unicode standard defines a system for numbering all of the common characters used in a large number of
languages. It is a suitable scheme for representing the information content of text, but not its appearance, since
Unicode values identify characters, not glyphs. For information about Unicode, see the Unicode Standard by
the Unicode Consortium (see the Bibliography).
When extracting character content, a conforming reader can easily convert text to Unicode values if a font’s
characters are identified according to a standard character set that is known to the conforming reader. This
character identification can occur if either the font uses a standard named encoding or the characters in the
font are identified by standard character names or CIDs in a well-known collection. 9.10.2, "Mapping Character
Codes to Unicode Values", describes in detail the overall algorithm for mapping character codes to Unicode
values.
If a font is not defined in one of these ways, the glyphs can still be shown, but the characters cannot be
converted to Unicode values without additional information:
This information can be provided as an optional ToUnicode entry in the font dictionary (PDF 1.2; see
9.10.3, "ToUnicode CMaps"), whose value shall be a stream object containing a special kind of CMap file
that maps character codes to Unicode values.
An ActualText entry for a structure element or marked-content sequence (see 14.9.4, "Replacement
Text") may be used to specify the text content directly.
9.10.2
Mapping Character Codes to Unicode Values
A conforming reader can use these methods, in the priority given, to map a character code to a Unicode value.
Tagged PDF documents, in particular, shall provide at least one of these methods (see 14.8.2.4.2, "Unicode
Mapping in Tagged PDF"):
If the font dictionary contains a ToUnicode CMap (see 9.10.3, "ToUnicode CMaps"), use that CMap to
convert the character code to Unicode.
If the font is a simple font that uses one of the predefined encodings MacRomanEncoding,
MacExpertEncoding, or WinAnsiEncoding, or that has an encoding whose Differences array includes
only character names taken from the Adobe standard Latin character set and the set of named characters
in the Symbol font (see Annex D):
a) Map the character code to a character name according to Table D.1 and the font’s Differences
array.
292
b) Look up the character name in the Adobe Glyph List (see the Bibliography) to obtain the
corresponding Unicode value.
If the font is a composite font that uses one of the predefined CMaps listed in Table 118 (except Identity-H
and Identity-V) or whose descendant CIDFont uses the Adobe-GB1, Adobe-CNS1, Adobe-Japan1, or
Adobe-Korea1 character collection:
a) Map the character code to a character identifier (CID) according to the font’s CMap.
b) Obtain the registry and ordering of the character collection used by the font’s CMap (for example,
Adobe and Japan1) from its CIDSystemInfo dictionary.
c) Construct a second CMap name by concatenating the registry and ordering obtained in step (b) in
the format registry-ordering-UCS2 (for example, Adobe-Japan1-UCS2).
d) Obtain the CMap with the name constructed in step (c) (available from the ASN Web site; see the
Bibliography).
e) Map the CID obtained in step (a) according to the CMap obtained in step (d), producing a
Unicode value.
NOTE
Type 0 fonts whose descendant CIDFonts use the Adobe-GB1, Adobe-CNS1, Adobe-Japan1, or Adobe-
Korea1 character collection (as specified in the CIDSystemInfo dictionary) shall have a supplement number
corresponding to the version of PDF supported by the conforming reader. See Table 3 for a list of the character
collections corresponding to a given PDF version. (Other supplements of these character collections can be
used, but if the supplement is higher-numbered than the one corresponding to the supported PDF version,
only the CIDs in the latter supplement are considered to be standard CIDs.)
If these methods fail to produce a Unicode value, there is no way to determine what the character code
represents in which case a conforming reader may choose a character code of their choosing.
9.10.3
ToUnicode CMaps
The CMap defined in the ToUnicode entry of the font dictionary shall follow the syntax for CMaps introduced in
9.7.5, "CMaps" and fully documented in Adobe Technical Note #5014, Adobe CMap and CIDFont Files
Specification. Additional guidance regarding the CMap defined in this entry is provided in Adobe Technical Note
#5411, ToUnicode Mapping File Tutorial. This CMap differs from an ordinary one in these ways:
The only pertinent entry in the CMap stream dictionary (see Table 120) is UseCMap, which may be used if
the CMap is based on another ToUnicode CMap.
The CMap file shall contain begincodespacerange and endcodespacerange operators that are
consistent with the encoding that the font uses. In particular, for a simple font, the codespace shall be one
byte long.
It shall use the beginbfchar, endbfchar, beginbfrange, and endbfrange operators to define the mapping
from character codes to Unicode character sequences expressed in UTF-16BE encoding.
EXAMPLE 1
This example illustrates a Type 0 font that uses the Identity-H CMap to map from character codes to CIDs
and whose descendant CIDFont uses the Identity mapping from CIDs to TrueType glyph indices. Text
strings shown using this font simply use a 2-byte glyph index for each glyph. In the absence of a
ToUnicode entry, no information would be available about what the glyphs mean.
14 0 obj
<<
/Type /Font
/Subtype /Type0
/BaseFont /Ryumin−Light
/Encoding /Identity−H
/DescendantFonts [ 15 0 R ]
/ToUnicode 16 0 R
293
>>
endobj
15 0 obj
<<
/Type /Font
/Subtype /CIDFontType2
/BaseFont /Ryumin−Light
/CIDSystemInfo 17 0 R
/FontDescriptor 18 0 R
/CIDToGIDMap /Identity
>>
endobj
EXAMPLE 2
In this example, the value of the ToUnicode entry is a stream object that contains the definition of the
CMap.
The begincodespacerange and endcodespacerange operators define the source character code range
to be the 2-byte character codes from < 00 00 > to < FF FF >. The specific mappings for several of the
character codes are shown.
16 0 obj
<< /Length 433 >>
stream
/CIDInit
/ProcSet findresource begin
12 dict begin
begincmap
/CIDSystemInfo
<< /Registry ( Adobe )
/Ordering ( UCS )
/Supplement 0
>> def
/CMapName /Adobe−Identity−UCS def
/CMapType 2 def
1 begincodespacerange
< 0000 >
< FFFF >
endcodespacerange
2 beginbfrange
< 0000 >
< 005E >
< 0020 >
< 005F >
< 0061 >
[<00660066 > < 00660069 > < 00660066006C > ]
endbfrange
1 beginbfchar
<3A51>
<D840DC3E>
endbfchar
endcmap
CMapName currentdict /CMap defineresource pop
end
end
endstream
endobj
< 00 00 > to < 00 5E > are mapped to the Unicode values U+0020 to U+007E This is followed by the
definition of a mapping where each character code represents more than one Unicode value:
< 005F > < 0061 > [ < 00660066 > < 00660069 > < 00660066006C > ]
In this case, the original character codes are the glyph indices for the ligatures ff, fi, and ffl. The entry
defines the mapping from the character codes < 00 5F >, < 00 60 >, and < 00 61 > to the strings of Unicode
values with a Unicode scalar value for each character in the ligature: U+0066 U+0066 are the Unicode
values for the character sequence f f, U+0066 U+0069 for f i, and U+0066 U+0066 U+006c for f f l.
Finally, the character code < 3A 51> is mapped to the Unicode value U+2003E, which is expressed by the
byte sequence <D840DC3E> in UTF-16BE encoding.
294
EXAMPLE 2 in this sub-clause illustrates several extensions to the way destination values may be defined. To
support mappings from a source code to a string of destination codes, this extension has been made to the
ranges defined after a beginbfchar operator:
n beginbfchar
srcCode dstString
endbfchar
where dstString may be a string of up to 512 bytes. Likewise, mappings after the beginbfrange operator may
be defined as:
n beginbfrange
srcCode1 srcCode2 dstString
endbfrange
In this case, the last byte of the string shall be incremented for each consecutive code in the source code
range.
When defining ranges of this type, the value of the last byte in the string shall be less than or equal to 255 −
(srcCode2 − srcCode1). This ensures that the last byte of the string shall not be incremented past 255;
otherwise, the result of mapping is undefined.
To support more compact representations of mappings from a range of source character codes to a
discontiguous range of destination codes, the CMaps used for the ToUnicode entry may use this syntax for the
mappings following a beginbfrange definition.
n beginbfrange
srcCode1 srcCode2 [ dstString1 dstString2dstStringm ]
endbfrange
Consecutive codes starting with srcCode1 and ending with srcCode2 shall be mapped to the destination strings
in the array starting with dstString1 and ending with dstStringm . The value of dstString can be a string of up to
512 bytes. The value of m represents the number of continuous character codes in the source character code
range.
m = srcCode2 - srcCode1 + 1
295
10
Rendering
10.1
General
Nearly all of the rendering facilities that are under the control of a PDF file pertain to the reproduction of colour.
Colours shall be rendered by a conforming reader using the following multiple-step process outlined.
NOTE 1
The PDF imaging model separates graphics (the specification of shapes and colours) from rendering
(controlling a raster output device). Figures 20 and 21 in 8.6.3, "Colour Space Families" illustrate this division.
8, "Graphics" describes the facilities for specifying the appearance of pages in a device-independent way. This
clause describes the facilities for controlling how shapes and colours are rendered on the raster output device.
All of the facilities discussed here depend on the specific characteristics of the output device. PDF files that are
intended to be device-independent should limit themselves to the general graphics facilities described in 8,
"Graphics".
Depending on the current colour space and on the characteristics of the device, it is not always necessary to
perform every step.
a) If a colour has been specified in a CIE-based colour space (see 8.6.5, "CIE-Based Colour Spaces"), it shall
first be transformed to the native colour space of the raster output device (also called its process colour
model).
b) If a colour has been specified in a device colour space that is inappropriate for the output device (for
example, RGB colour with a CMYK or grayscale device), a colour conversion function shall be invoked.
c) The device colour values shall now be mapped through transfer functions, one for each colour component.
NOTE 2
The transfer functions compensate for peculiarities of the output device, such as nonlinear gray-level
response. This step is sometimes called gamma correction.
d) If the device cannot reproduce continuous tones, but only certain discrete colours such as black and white
pixels, a halftone function shall be invoked, which approximates the desired colours by means of patterns
of pixels.
e) Finally, scan conversion shall be performed to mark the appropriate pixels of the raster output device with
the requested colours.
Once these operations have been performed for all graphics objects on the page, the resulting raster data shall
be used to mark the physical output medium, such as pixels on a display or ink on a printed page. A PDF file
may specify very little about the properties of the physical medium on which the output will be produced; that
information may be obtained from the following sources by a conforming reader:
The media box and a few other entries in the page dictionary (see 14.11.2, "Page Boundaries").
An interactive dialogue conducted when the user requests viewing or printing.
A job ticket, either embedded in the PDF file or provided separately, that may specify detailed instructions
for imposing PDF pages onto media and for controlling special features of the output device. Various
standards exist for the format of job tickets. Two of them, JDF (Job Definition Format) and PJTF (Portable
Job Ticket Format), are described in the CIP4 document JDF Specification and in Adobe Technical Note
#5620, Portable Job Ticket Format (see the Bibliography), respectively.
Table 58 in 8.4.5, "Graphics State Parameter Dictionaries" lists the various device-dependent graphics state
parameters that may be used to control certain aspects of rendering. To invoke these parameters, the gs
operator shall be used.
296
10.2
CIE-Based Colour to Device Colour
To render CIE-based colours on an output device, the conforming reader shall convert from the specified CIE-
based colour space to the device’s native colour space (typically DeviceGray, DeviceRGB, or DeviceCMYK),
taking into account the known properties of the device.
NOTE 1
As discussed in 8.6.5, "CIE-Based Colour Spaces" CIE-based colour is based on a model of human colour
perception. The goal of CIE-based colour rendering is to produce output in the device’s native colour space
that accurately reproduces the requested CIE-based colour values as perceived by a human observer. CIE-
based colour specification and rendering are a feature of PDF 1.1 (CalGray, CalRGB, and Lab) and PDF 1.3
(ICCBased).
NOTE 2
The conversion from CIE-based colour to device colour is complex, and the theory on which it is based is
beyond the scope of this specification. The algorithm has many parameters, including an optional, full three-
dimensional colour lookup table. The colour fidelity of the output depends on having these parameters properly
set, usually by a method that includes some form of calibration. The colours that a device can produce are
characterized by a device profile, which is usually specified by an ICC profile associated with the device (and
entirely separate from the profile that is specified in an ICCBased colour space).
NOTE 3
PDF has no equivalent of the PostScript colour rendering dictionary. The means by which a device profile is
associated with a conforming reader’s output device are implementation-dependent and not specified in a PDF
file. Typically, this is done through a colour management system (CMS) that is provided by the operating
system. Beginning with PDF 1.4, a PDF file can also specify one or more output intents providing possible
profiles that may be used to process the file (see 14.11.5, "Output Intents").
Conversion from a CIE-based colour value to a device colour value requires two main operations:
a) The CIE-based colour value shall be adjusted according to a CIE-based gamut mapping function.
NOTE 4
A gamut is a subset of all possible colours in some colour space. A page description has a source gamut
consisting of all the colours it uses. An output device has a device gamut consisting of all the colours it can
reproduce. This step transforms colours from the source gamut to the device gamut in a way that attempts to
preserve colour appearance, visual contrast, or some other explicitly specified rendering intent (see 8.6.5.8,
"Rendering Intents").
b) A corresponding device colour value shall be generated according to a CIE-based colour mapping
function. For a given CIE-based colour value, this function shall compute a colour value in the device’s
native colour space.
The CIE-based gamut and colour mapping functions shall be applied only to colour values presented in a CIE-
based colour space. Colour values in device colour spaces directly control the device colour components
though this may be altered by the DefaultGray, DefaultRGB, and DefaultCMYK colour space resources (see
8.6.5.6, "Default Colour Spaces").
The source gamut shall be specified by the information contained in the definition of the CIE-based colour
space when selected. This specification shall be device-independent. The corresponding properties of the
output device shall be given in the device profile associated with the device. The gamut mapping and colour
mapping functions are part of the implementation of the conforming reader.
10.3
Conversions among Device Colour Spaces
10.3.1
General
Each raster output device has a native colour space, which typically is one of the standard device colour
spaces (DeviceGray, DeviceRGB, or DeviceCMYK). In other words, most devices support reproduction of
colours according to a grayscale (monochrome), RGB (red-green-blue), or CMYK (cyan-magenta-yellow-black)
model. If the device supports continuous-tone output, reproduction shall occur directly. Otherwise, it shall be
accomplished by means of halftoning.
297
A device’s native colour space is also called its process colour model. Process colours are ones that are
produced by combinations of one or more standard process colorants. Colours specified in any device or CIE-
based colour space shall be rendered as process colours. A device may also support additional spot colorants,
which shall be painted only by means of Separation or DeviceN colour spaces. They shall not be involved in
the rendering of device or CIE-based colour spaces, nor shall they be subject to the conversions described in
the Note.
NOTE
Some devices provide a native colour space that is not one of the three named previously but consists of a
different combination of colorants. In that case, conversion from the standard device colour spaces to the
device’s native colour space may be performed by the conforming reader in a manner of its own choosing.
Knowing the native colour space and other output capabilities of the device, the conforming reader shall
automatically convert the colour values specified in a file to those appropriate for the device’s native colour
space. If the file specifies colours directly in the device’s native colour space, no conversions shall be
performed.
EXAMPLE
If a file specifies colours in the DeviceRGB colour space but the device supports grayscale (such as a
monochrome display) or CMYK (such as a colour printer), the conforming reader shall perform the
necessary conversions.
The algorithms used to convert among device colour spaces are very simple. As perceived by a human viewer,
these conversions produce only crude approximations of the original colours. More sophisticated control over
colour conversion may be achieved by means of CIE-based colour specification and rendering. Additionally,
device colour spaces may be remapped into CIE-based colour spaces (see 8.6.5.6, "Default Colour Spaces").
10.3.2
Conversion between DeviceGray and DeviceRGB
Black, white, and intermediate shades of gray can be considered special cases of RGB colour. A grayscale
value shall be described by a single number: 0.0 corresponds to black, 1.0 to white, and intermediate values to
different gray levels.
A gray level shall be equivalent to an RGB value with all three components the same. In other words, the RGB
colour value equivalent to a specific gray value shall be
red = gray
green = gray
blue = gray
The gray value for a given RGB value shall be computed according to the NTSC video standard, which
determines how a colour television signal is rendered on a black-and-white television set:
gray = 0.3 × red + 0.59 × green + 0.11 × blu
10.3.3
Conversion between DeviceGray and DeviceCMYK
Nominally, a gray level is the complement of the black component of CMYK. Therefore, the CMYK colour value
equivalent to a specific gray level shall be
cyan = 0.0
magenta = 0.0
yellow = 0.0
black = 1.0 - gray
To obtain the equivalent gray level for a given CMYK value, the contributions of all components shall be taken
into account:
·
gray
=
1.0
min(1.0, 0.3 × cyan + 0.59 × magenta + 0.11 × yellow + black
298
The interactions between the black component and the other three are elaborated in 10.3.4.
10.3.4
Conversion from DeviceRGB to DeviceCMYK
Conversion of a colour value from RGB to CMYK is a two-step process. The first step shall be to convert the
red-green-blue value to equivalent cyan, magenta, and yellow components. The second step shall be to
generate a black component and alter the other components to produce a better approximation of the original
colour.
NOTE 1
The subtractive colour primaries cyan, magenta, and yellow are the complements of the additive primaries red,
green, and blue.
EXAMPLE
A cyan ink subtracts the red component of white light. In theory, the conversion is very simple:
cyan = 1.0 - red
magenta = 1.0 - green
yellow = 1.0 - blue
A colour that is 0.2 red, 0.7 green, and 0.4 blue can also be expressed as 1.0 − 0.2 = 0.8 cyan, 1.0 −
0.7 = 0.3 magenta, and 1.0 − 0.4 = 0.6 yellow.
NOTE 2
Logically, only cyan, magenta, and yellow are needed to generate a printing colour. An equal level of cyan,
magenta, and yellow should create the equivalent level of black. In practice, however, coloured printing inks do
not mix perfectly; such combinations often form dark brown shades instead of true black. To obtain a truer
colour rendition on a printer, true black ink is often substituted for the mixed-black portion of a colour. Most
colour printers support a black component (the K component of CMYK). Computing the quantity of this
component requires some additional steps:
Black generation calculates the amount of black to be used when trying to reproduce a particular colour.
Undercolor removal reduces the amounts of the cyan, magenta, and yellow components to compensate for the
amount of black that was added by black generation.
The complete conversion from RGB to CMYK shall be as follows, where BG (k) and UCR (k) are invocations of
the black-generation and undercolor-removal functions, respectively:
c = 1.0 - red
m = 1.0 - green
y = 1.0 - blue
k = min(c, m, y)
cyan = min (1.0, max (0.0, c - UCR (k)))
magenta = min (1.0, max (0.0, m - UCR (k)))
yellow = min (1.0, max (0.0, y - UCR(k)))
black = min (1.0, max (0.0, BG (k)))
The black-generation and undercolor-removal functions shall be defined as PDF function dictionaries (see
7.10, "Functions") that are parameters in the graphics state. They shall be specified as the values of the BG
and UCR (or BG2 and UCR2) entries in a graphics state parameter dictionary (see Table 58). Each function
shall be called with a single numeric operand and shall return a single numeric result.
The input of both the black-generation and undercolor-removal functions shall be k, the minimum of the
intermediate c, m, and y values that have been computed by subtracting the original red, green, and blue
components from 1.0.
NOTE 3
Nominally, k is the amount of black that can be removed from the cyan, magenta, and yellow components and
substituted as a separate black component.
299
The black-generation function shall compute the black component as a function of the nominal k value. It may
simply return its k operand unchanged, or it may return a larger value for extra black, a smaller value for less
black, or 0.0 for no black at all.
The undercolor-removal function shall compute the amount to subtract from each of the intermediate c, m, and
y values to produce the final cyan, magenta, and yellow components. It may simply return its k operand
unchanged, or it may return 0.0 (so that no colour is removed), some fraction of the black amount, or even a
negative amount, thereby adding to the total amount of colorant.
The final component values that result after applying black generation and undercolor removal should be in the
range 0.0 to 1.0. If a value falls outside this range, the nearest valid value shall be substituted automatically
without error indication.
NOTE 4
This substitution is indicated explicitly by the min and max operations in the preceding formulas.
The correct choice of black-generation and undercolor-removal functions depends on the characteristics of the
output device. Each device shall be configured with default values that are appropriate for that device.
NOTE 5
See 11.7.5, "Rendering Parameters and Transparency" and, in particular, 11.7.5.3, "Rendering Intent and
Colour Conversions" for further discussion of the role of black-generation and undercolor-removal functions in
the transparent imaging model.
10.3.5
Conversion from DeviceCMYK to DeviceRGB
Conversion of a colour value from CMYK to RGB is a simple operation that does not involve black generation
or undercolour removal:
red = 1.0 - min(1.0, cyan + black)
green = 1.0 - min(1.0, magenta + black)
blue = 1.0 - min(1.0, yellow + black)
The black component shall be added to each of the other components, which shall then be converted to their
complementary colours by subtracting them each from 1.0.
10.4
Transfer Functions
In the sequence of steps for processing colours, the conforming reader shall apply the transfer function after
performing any needed conversions between colour spaces, but before applying a halftone function, if
necessary. Each colour component shall have its own separate transfer function; there shall not be interaction
between components.
NOTE 1
Starting with PDF 1.2, a transfer function may be used to adjust the values of colour components to
compensate for nonlinear response in an output device and in the human eye. Each component of a device
colour space—for example, the red component of the DeviceRGB space—is intended to represent the
perceived lightness or intensity of that colour component in proportion to the component’s numeric value.
NOTE 2
Many devices do not actually behave this way, however; the purpose of a transfer function is to compensate for
the device’s actual behaviour. This operation is sometimes called gamma correction (not to be confused with
the CIE-based gamut mapping function performed as part of CIE-based colour rendering).
Transfer functions shall always operate in the native colour space of the output device, regardless of the colour
space in which colours were originally specified. (For example, for a CMYK device, the transfer functions apply
to the device’s cyan, magenta, yellow, and black colour components, even if the colours were originally
specified in, for example, a DeviceRGB or CalRGB colour space.) The transfer function shall be called with a
numeric operand in the range 0.0 to 1.0 and shall return a number in the same range. The input shall be the
value of a colour component in the device’s native colour space, either specified directly or produced by
conversion from some other colour space. The output shall be the transformed component value to be
transmitted to the device (after halftoning, if necessary).
300
Both the input and the output of a transfer function shall always be interpreted as if the corresponding colour
component were additive (red, green, blue, or gray): the greater the numeric value, the lighter the colour. If the
component is subtractive (cyan, magenta, yellow, black, or a spot colour), it shall be converted to additive form
by subtracting it from 1.0 before it is passed to the transfer function. The output of the function shall always be
in additive form and shall be passed on to the halftone function in that form.
Starting with PDF 1.2, transfer functions shall be defined as PDF function objects (see 7.10, "Functions").
There are two ways to specify transfer functions:
The current transfer function parameter in the graphics state shall consist of either a single transfer
function or an array of four separate transfer functions, one each for red, green, blue, and gray or their
complements cyan, magenta, yellow, and black. If only a single function is specified, it shall apply to all
components. An RGB device shall use the first three, a monochrome device shall use the gray transfer
function only, and a CMYK device shall use all four. The current transfer function may be specified as the
value of the TR or TR2 entry in a graphics state parameter dictionary; see Table 58.
The current halftone parameter in the graphics state may specify transfer functions as optional entries in
halftone dictionaries (see 10.5.5, "Halftone Dictionaries"). This is the only way to set transfer functions for
nonprimary colour components or for any component in devices whose native colour space uses
components other than the ones listed previously. A transfer function specified in a halftone dictionary shall
override the corresponding one specified by the current transfer function parameter in the graphics state.
In addition to their intended use for gamma correction, transfer functions may be used to produce a variety of
special, device-dependent effects. Because transfer functions produce device-dependent effects, a page
description that is intended to be device-independent shall not alter them.
When the current colour space is DeviceGray and the output device’s native colour space is DeviceCMYK, a
conforming reader shall use only the gray transfer function. The normal conversion from DeviceGray to
DeviceCMYK produces 0.0 for the cyan, magenta, and yellow components. These components shall not be
passed through their respective transfer functions but are rendered directly, producing output containing no
coloured inks. This special case exists for compatibility with existing conforming readers that use a transfer
function to obtain special effects on monochrome devices, and shall apply only to colours specified in the
DeviceGray colour space.
NOTE 3
See 11.7.5, "Rendering Parameters and Transparency" and, in particular, 11.7.5.2, "Halftone and Transfer
Function" for further discussion of the role of transfer functions in the transparent imaging model.
10.5
Halftones
10.5.1
General
Halftoning is a process by which continuous-tone colours are approximated on an output device that can
achieve only a limited number of discrete colours. Colours that the device cannot produce directly are
simulated by using patterns of pixels in the colours available.
NOTE 1
Perhaps the most familiar example is the rendering of gray tones with black and white pixels, as in a
newspaper photograph.
Some output devices can reproduce continuous-tone colours directly. Halftoning is not required for such
devices; after gamma correction by the transfer functions, the colour components shall be transmitted directly
to the device. On devices that do require halftoning, it shall occur after all colour components have been
transformed by the applicable transfer functions. The input to the halftone function shall consist of continuous-
tone, gamma-corrected colour components in the device’s native colour space. Its output shall consist of pixels
in colours the device can reproduce.
PDF provides a high degree of control over details of the halftoning process.
NOTE 2
When rendering on low-resolution displays, fine control over halftone patterns is needed to achieve the best
approximations of gray levels or colours and to minimize visual artifacts.
301
NOTE 3
In colour printing, independent halftone screens can be specified for each of several colorants.
NOTE 4
Remember that everything pertaining to halftones is, by definition, device-dependent. In general, when a PDF
file provides its own halftone specifications, it sacrifices portability. Associated with every output device is a
default halftone definition that is appropriate for most purposes. Only relatively sophisticated files need to
define their own halftones to achieve special effects. For correct results, a PDF file that defines a new halftone
depends on certain assumptions about the resolution and orientation of device space. The best choice of
halftone parameters often depends on specific physical properties of the output device, such as pixel shape,
overlap between pixels, and the effects of electronic or mechanical noise.
All halftones are defined in device space, and shall be unaffected by the current transformation matrix.
10.5.2
Halftone Screens
In general, halftoning methods are based on the notion of a halftone screen, which divides the array of device
pixels into cells that may be modified to produce the desired halftone effects. A screen is defined by
conceptually laying a uniform rectangular grid over the device pixel array. Each pixel belongs to one cell of the
grid; a single cell typically contains many pixels. The screen grid shall be defined entirely in device space and
shall be unaffected by modifications to the current transformation matrix.
NOTE
This property is essential to ensure that adjacent areas coloured by halftones are properly stitched together
without visible seams.
On a bilevel (black-and-white) device, each cell of a screen may be made to approximate a shade of gray by
painting some of the cell’s pixels black and some white. Numerically, the gray level produced within a cell shall
be the ratio of white pixels to the total number of pixels in the cell. A cell containing n pixels can render n + 1
different gray levels, ranging from all pixels black to all pixels white. A gray value g in the range 0.0 to 1.0 shall
be produced by making i pixels white, where i = floor (g × n).
The foregoing description also applies to colour output devices whose pixels consist of primary colours that are
either completely on or completely off. Most colour printers, but not colour displays, work this way. Halftoning
shall be applied to each colour component independently, producing shades of that colour.
Colour components shall be presented to the halftoning machinery in additive form, regardless of whether they
were originally specified additively (RGB or gray) or subtractively (CMYK or tint). Larger values of a colour
component represent lighter colours—greater intensity in an additive device such as a display or less ink in a
subtractive device such as a printer. Transfer functions produce colour values in additive form; see 10.4,
"Transfer Functions".
10.5.3
Spot Functions
A common way of defining a halftone screen is by specifying a frequency, angle, and spot function. The
frequency indicates the number of halftone cells per inch; the angle indicates the orientation of the grid lines
relative to the device coordinate system. As a cell’s desired gray level varies from black to white, individual
pixels within the cell change from black to white in a well-defined sequence: if a particular gray level includes
certain white pixels, lighter grays will include the same white pixels along with some additional ones. The order
in which pixels change from black to white for increasing gray levels shall be determined by a spot function,
which specifies that order in an indirect way that minimizes interactions with the screen frequency and angle.
Consider a halftone cell to have its own coordinate system: the centre of the cell is the origin and the corners
are at coordinates ±1.0 horizontally and vertically. Each pixel in the cell is centred at horizontal and vertical
coordinates that both lie in the range −1.0 to +1.0. For each pixel, the spot function shall be invoked with the
pixel’s coordinates as input and shall return a single number in the range −1.0 to +1.0, defining the pixel’s
position in the whitening order.
The specific values the spot function returns are not significant; all that matters are the relative values returned
for different pixels. As a cell’s gray level varies from black to white, the first pixel whitened shall be the one for
which the spot function returns the lowest value, the next pixel shall be the one with the next higher spot
302
function value, and so on. If two pixels have the same spot function value, their relative order shall be chosen
arbitrarily.
PDF provides built-in definitions for many of the most commonly used spot functions that a conforming reader
shall implement. A halftone may simply specify any of these predefined spot functions by name instead of
giving an explicit function definition.
EXAMPLE
The name SimpleDot designates a spot function whose value is inversely related to a pixel’s distance
from the center of the halftone cell. This produces a “dot screen” in which the black pixels are clustered
within a circle whose area is inversely proportional to the gray level. The name Line designates a spot
function whose value is the distance from a given pixel to a line through the center of the cell, producing a
“line screen” in which the white pixels grow away from that line.
Table 128 shows the predefined spot functions. The table gives the mathematical definition of each function
along with the corresponding PostScript language code as it would be defined in a PostScript calculator
function (see 7.10.5, "Type 4 (PostScript Calculator) Functions"). The image accompanying each function
shows how the relative values of the function are distributed over the halftone cell, indicating the approximate
order in which pixels are whitened. Pixels corresponding to darker points in the image are whitened later than
those corresponding to lighter points.
Table 128 - Predefined spot functions
Name
Appearance
Definition
SimpleDot
1 − (x 2 + y 2 )
{ dup mul exch dup mul add 1 exch sub }
InvertedSimpleDot
x 2 + y 2 − 1
{ dup mul exch dup mul add 1 sub }
DoubleDot
sin (360 × x)
+ ------------------------------
2
2
{
360 mul sin 2 div exch 360 mul sin 2 div add }
InvertedDoubleDot
sin(360 × x)
+ ------------------------------
2
2
{
360 mul sin 2 div exch 360 mul sin 2 div add neg }
303
Table 128 - Predefined spot functions (continued)
Name
Appearance
Definition
CosineDot
cos (180 × x)
+ -------------------------------
2
2
{
180 mul cos exch 180 mul cos add 2 div }
Double
sin
360
×
-
2
-------------------------------
+ ------------------------------
2
2
{
360 mul sin 2 div exch 2 div 360 mul sin 2 div add }
InvertedDouble
x
sin
360
× --
2
-
-------------------------------
+ ------------------------------
2
2
{
360 mul sin 2 div exch
2 div 360 mul sin 2
div add neg }
Line
−| y|
{ exch pop abs neg }
LineX
x
{ pop }
LineY
y
{ exch pop }
304
Table 128 - Predefined spot functions (continued)
Name
Appearance
Definition
Round
if |x| + |y|
≤ 1 then 1 − (x 2 + y 2)
else (| x| − 1) 2 + (| y| − 1) 2 − 1
{ abs exch abs
2 copy add 1 le
{ dup mul exch dup mul add 1 exch sub }
{
1 sub dup mul exch 1 sub dup mul add 1
sub }
ifelse
}
Ellipse
let w = (3 × | x| ) + (4 × | y|) − 3
2
y
2
x
+
0.75
if w < 0 then
1
- ------------------------------
4
2
1- y
2
(1- x )
+
0.75
else if w > 1 then
----------------------------------------------------- - 1
4
else 0.5 − w
{ abs exch abs 2 copy
3 mul exch
4 mul add 3
sub dup 0 lt
{ pop dup mul exch 0.75 div dup mul add
4 div 1 exch sub }
{ dup 1 gt
{ pop 1 exch sub dup mul
exch 1 exch sub 0.75 div dup
mul add
4 div 1 sub }
{
0.5 exch sub exch pop exch pop }
ifelse
}
ifelse
}
EllipseA
1 − (x 2 + 0.9 × y 2 )
{ dup mul 0.9 mul exch dup mul add 1 exch sub }
InvertedEllipseA
x 2 + 0.9 × y 2 − 1
{ dup mul 0.9 mul exch dup mul add 1 sub }
305
Table 128 - Predefined spot functions (continued)
Name
Appearance
Definition
EllipseB
1
-
x2
+
-
×
y2
8
{ dup 5 mul 8 div mul exch dup mul exch add sqrt
1 exch sub }
EllipseC
1 − (0.9 × x2 + y2 )
{ dup mul exch dup mul 0.9 mul add 1 exch sub }
InvertedEllipseC
0.9 × x 2 + y 2 − 1
{ dup mul exch dup mul 0.9 mul add 1 sub }
Square
−max (| x| , | y| )
{ abs exch abs 2 copy lt
{ exch }
if
pop neg }
Cross
−min (| x| , | y| )
{ abs exch abs 2 copy gt
{ exch }
if
pop neg }
Rhomboid
0.9 × x
+ y
2
{ abs exch abs 0.9 mul add 2 div }
306
Table 128 - Predefined spot functions (continued)
Name
Appearance
Definition
Diamond
if |x| + |y| ≤ 0.75 then 1 − (x 2 + y 2)
else if | x| + | y| ≤ 1.23 then 1 − (0.85 × | x | + | y | )
else ( | x | − 1) 2 + ( | y | − 1) 2 − 1
{ abs exch abs 2 copy add 0.75 le
{ dup mul exch dup mul add 1 exch sub }
{
2 copy add 1.23 le
{
0.85 mul add 1 exch sub }
{
1 sub dup mul exch 1 sub du
mul add 1 sub }
ifelse
}
ifelse
}
Figure 49 illustrates the effects of some of the predefined spot functions.
150 per inch at 45
100 per inch at 45
50 per inch at 45
75 per inch at 45
Round dot screen Round dot screen Round dot screen
Line screen
Figure 49 - Various halftoning effects
10.5.4
Threshold Arrays
Another way to define a halftone screen is with a threshold array that directly controls individual device pixels in
a halftone cell. This technique provides a high degree of control over halftone rendering. It also permits halftone
cells to be arbitrary rectangles, whereas those controlled by a spot function are always square.
A threshold array is much like a sampled image—a rectangular array of pixel values—but shall be defined
entirely in device space. Depending on the halftone type, the threshold values occupy 8 or 16 bits each.
Threshold values nominally represent gray levels in the usual way, from 0 for black up to the maximum (255 or
65,535) for white. The threshold array shall be replicated to tile the entire device space: each pixel in device
space shall be mapped to a particular sample in the threshold array. On a bilevel device, where each pixel is
either black or white, halftoning with a threshold array shall proceed as follows:
a) For each device pixel that is to be painted with some gray level, consult the corresponding threshold value
from the threshold array.
b) If the requested gray level is less than the threshold value, paint the device pixel black; otherwise, paint it
white. Gray levels in the range 0.0 to 1.0 correspond to threshold values from 0 to the maximum available
(255 or 65,535).
307
A threshold value of 0 shall be treated as if it were 1; therefore, a gray level of 0.0 paints all pixels black,
regardless of the values in the threshold array.
This scheme easily generalizes to monochrome devices with multiple bits per pixel, where each pixel can
directly represent intermediate gray levels in addition to black and white. For any device pixel that is specified
with some in-between gray level, the halftoning algorithm shall consult the corresponding value in the threshold
array to determine whether to use the next-lower or next-higher representable gray level. In this situation, the
threshold values do not represent absolute gray levels, but rather gradations between any two adjacent
representable gray levels.
EXAMPLE
If there are 2 bits per pixel, each pixel can directly represent one of four different gray levels: black, dark
gray, light gray, or white, encoded as 0, 1, 2, and 3, respectively.
NOTE
A halftone defined in this way can also be used with colour displays that have a limited number of values for
each colour component. The red, green, and blue components are simply treated independently as gray
levels, applying the appropriate threshold array to each. (This technique also works for a screen defined as a
spot function, since the spot function is used to compute a threshold array internally.)
10.5.5
Halftone Dictionaries
10.5.5.1
General
In PDF 1.2, the graphics state includes a current halftone parameter, which determines the halftoning process
that a conforming reader shall use to perform painting operations. The current halftone may be specified as the
value of the HT entry in a graphics state parameter dictionary; see Table 58. It may be defined by either a
dictionary or a stream, depending on the type of halftone; the term halftone dictionary is used generically
throughout this clause to refer to either a dictionary object or the dictionary portion of a stream object. (The
halftones that are defined by streams are specifically identified as such in the descriptions of particular halftone
types; unless otherwise stated, they are understood to be defined by simple dictionaries instead.)
Every halftone dictionary shall have a HalftoneType entry whose value shall be an integer specifying the
overall type of halftone definition. The remaining entries in the dictionary are interpreted according to this type.
PDF supports the halftone types listed in Table 129.
Table 129 - PDF halftone types
Type
Meaning
1
Defines a single halftone screen by a frequency, angle, and spot
function.
5
Defines an arbitrary number of halftone screens, one for each colorant
or colour component (including both primary and spot colorants). The
keys in this dictionary are names of colorants; the values are halftone
dictionaries of other types, each defining the halftone screen for a single
colorant.
6
Defines a single halftone screen by a threshold array containing 8-bit
sample values.
10
Defines a single halftone screen by a threshold array containing 8-bit
sample values, representing a halftone cell that may have a nonzero
screen angle.
16
(PDF 1.3) Defines a single halftone screen by a threshold array
containing 16-bit sample values, representing a halftone cell that may
have a nonzero screen angle.
NOTE 1
The dictionaries representing these halftone types contain the same entries as the corresponding PostScript
language halftone dictionaries (as described in Section 7.4 of the PostScript Language Reference, Third
Edition), with the following exceptions:
308
The PDF dictionaries may contain a Type entry with the value Halftone, identifying the type of PDF object that
the dictionary describes.
Spot functions and transfer functions are represented by function objects instead of PostScript procedures.
Threshold arrays are specified as streams instead of files.
In type 5 halftone dictionaries, the keys for colorants shall be name objects; they may not be strings as they
may in PostScript.
Halftone dictionaries have an optional entry, HalftoneName, that identifies the halftone by name. In PDF 1.3, if
this entry is present, all other entries, including HalftoneType, are optional. At rendering time, if the output
device has a halftone with the specified name, that halftone shall be used, overriding any other halftone
parameters specified in the dictionary.
NOTE 2
This provides a way for PDF files to select the proprietary halftones supplied by some device manufacturers,
which would not otherwise be accessible because they are not explicitly defined in PDF.
If there is no HalftoneName entry, or if the requested halftone name does not exist on the device, the halftone’s
parameters shall be defined by the other entries in the dictionary, if any. If no other entries are present, the
default halftone shall be used.
NOTE 3
See 11.7.5, "Rendering Parameters and Transparency" and, in particular, “Halftone and Transfer Function” in
11.7.5.2 for further discussion of the role of halftones in the transparent imaging model.
10.5.5.2
Type 1 Halftones
Table 130 describes the contents of a halftone dictionary of type 1, which defines a halftone screen in terms of
its frequency, angle, and spot function.
Table 130 - Entries in a type 1 halftone dictionary
Key
Type
Value
Type
name
(Optional) The type of PDF object that this dictionary describes;
if present, shall be Halftone for a halftone dictionary.
HalftoneType
integer
(Required) A code identifying the halftone type that this
dictionary describes; shall be 1 for this type of halftone.
HalftoneName
byte string
(Optional) The name of the halftone dictionary.
Frequency
number
(Required) The screen frequency, measured in halftone cells per
inch in device space.
Angle
number
(Required) The screen angle, in degrees of rotation
counterclockwise with respect to the device coordinate system.
NOTE
Most output devices have left-handed device
spaces. On such devices, a counterclockwise angle
in device space corresponds to a clockwise angle
in default user space and on the physical medium.
SpotFunction
function or name
(Required) A function object defining the order in which device
pixels within a screen cell shall be adjusted for different gray
levels, or the name of one of the predefined spot functions (see
Table 128).
AccurateScreens
boolean
(Optional) A flag specifying whether to invoke a special halftone
algorithm that is extremely precise but computationally
expensive; see Note 1 for further discussion. Default value:
false.
309
Table 130 - Entries in a type 1 halftone dictionary (continued)
Key
Type
Value
TransferFunction
function or name
(Optional) A transfer function, which overrides the current
transfer function in the graphics state for the same component.
This entry shall be present if the dictionary is a component of a
type
5 halftone
(see
“Type
5 Halftones” in
10.5.5.6) and
represents either a nonprimary or nonstandard primary colour
component (see 10.4, "Transfer Functions"). The name Identity
may be used to specify the identity function.
If the AccurateScreens entry has a value of true, a highly precise halftoning algorithm shall be substituted in
place of the standard one. If AccurateScreens is false or not present, ordinary halftoning shall be used.
NOTE 1
Accurate halftoning achieves the requested screen frequency and angle with very high accuracy, whereas
ordinary halftoning adjusts them so that a single screen cell is quantized to device pixels. High accuracy is
important mainly for making colour separations on high-resolution devices. However, it may be computationally
expensive and therefore is ordinarily disabled.
NOTE 2
In principle, PDF permits the use of halftone screens with arbitrarily large cells—in other words, arbitrarily low
frequencies. However, cells that are very large relative to the device resolution or that are oriented at
unfavorable angles may exceed the capacity of available memory. If this happens, an error occurs. The
AccurateScreens feature often requires very large amounts of memory to achieve the highest accuracy.
EXAMPLE
The following shows a halftone dictionary for a type 1 halftone.
28 0 obj
<< /Type /Halftone
/HalftoneType 1
/Frequency 120
/Angle 30
/SpotFunction /CosineDot
/TransferFunction /Identity
>>
endobj
10.5.5.3
Type 6 Halftones
A type 6 halftone defines a halftone screen with a threshold array. The halftone shall be represented as a
stream containing the threshold values; the parameters defining the halftone shall be specified by entries in the
stream dictionary. This dictionary may contain the entries shown in Table 131 in addition to the usual entries
common to all streams (see Table 5). The Width and Height entries shall specify the dimensions of the
threshold array in device pixels; the stream shall contain Width × Height bytes, each representing a single
threshold value. Threshold values are defined in device space in the same order as image samples in image
space (see Figure 34), with the first value at device coordinates (0, 0) and horizontal coordinates changing
faster than vertical coordinates.
Table 131 - Additional entries specific to a type 6 halftone dictionary
Key
Type
Value
Type
name
(Optional) The type of PDF object that this dictionary describes;
if present, shall be Halftone for a halftone dictionary.
HalftoneType
integer
(Required) A code identifying the halftone type that this
dictionary describes; shall be 6 for this type of halftone.
HalftoneName
byte string
(Optional) The name of the halftone dictionary.
Width
integer
(Required) The width of the threshold array, in device pixels.
Height
integer
(Required) The height of the threshold array, in device pixels.
310
Table 131 - Additional entries specific to a type 6 halftone dictionary (continued)
Key
Type
Value
TransferFunction
function or name
(Optional) A transfer function, which shall override the current
transfer function in the graphics state for the same component.
This entry shall be present if the dictionary is a component of a
type
5 halftone
(see
“Type
5 Halftones” in
10.5.5.6) and
represents either a nonprimary or nonstandard primary colour
component (see 10.4, "Transfer Functions"). The name Identity
may be used to specify the identity function.
10.5.5.4
Type 10 Halftones
Type 6 halftones specify a threshold array with a zero screen angle; they make no provision for other angles.
The type 10 halftone removes this restriction and allows the use of threshold arrays for halftones with nonzero
screen angles as well.
Halftone cells at nonzero angles can be difficult to specify because they may not line up well with scan lines
and because it may be difficult to determine where a given sampled point goes. The type 10 halftone addresses
these difficulties by dividing the halftone cell into a pair of squares that line up at zero angles with the output
device’s pixel grid. The squares contain the same information as the original cell but are much easier to store
and manipulate. In addition, they can be mapped easily into the internal representation used for all rendering.
NOTE 1
Figure 50 shows a halftone cell with a frequency of 38.4 cells per inch and an angle of 50.2 degrees,
represented graphically in device space at a resolution of 300 dots per inch. Each asterisk in the figure
represents a location in device space that is mapped to a specific location in the threshold array.
Figure 50 - Halftone cell with a nonzero angle
NOTE 2
Figure 51 shows how the halftone cell can be divided into two squares. If the squares and the original cell are
tiled across device space, the area to the right of the upper square maps exactly into the empty area of the
lower square, and vice versa (see Figure 52). The last row in the first square is immediately adjacent to the first
row in the second square and starts in the same column.
311
X
Y
*
Figure 51 - Angled halftone cell divided into two squares
Y
X
X
a
a
a a
a a
a a a
a a a
a a a a
a a a a
a a a a a
a a a a a
Y
a a a a a a
a a a a
b
b
a a a a a
a a a
b b
b b
a a a a
a a
b b b
b b b
a a a
a
b b b b
b b b b
a a
b b b b b
b b b
* b
Y
a
b b b b b b
b b b b
c
c b b b b b
b b b
X
c c
c c b b b b
b b
c c c
c c c b b b
b
c c c c
c c c c
b b
c c c c c
c c c
b
* c
c c c c c c
c c c c
c c c c c
c c c
c c c c
c c
c c c
c
c c
Y
c
Figure 52 - Halftone cell and two squares tiled across device space
NOTE 3
Any halftone cell can be divided in this way. The side of the upper square (X) is equal to the horizontal
displacement from a point in one halftone cell to the corresponding point in the adjacent cell, such as those
marked by asterisks in Figure 52. The side of the lower square (Y) is the vertical displacement between the
same two points. The frequency of a halftone screen constructed from squares with sides X and Y is thus
given by
frequency
= -----------------------
X2
+
Y2
and the angle by
angle
= atan
--
X
312

 

 

 

 

 

 

 

Content      ..     6      7      8      9     ..