Configure clients for Unicode
In a Unicode-mode environment, P4 Server clients use character set settings to translate text and metadata correctly between the server and the workstation. For example, a UNIX client examines the LANG or
LOCALE variables to determine the appropriate character set. In most cases, the correct character set is detected automatically. However, some situations require additional configuration or manual settings.
| Situation | Action |
|---|---|
|
The automatically selected character set is producing incorrect translations or display issues. |
See Troubleshoot user workstations in Unicode installations for more information about common Unicode-related workstation issues. |
|
On Windows workstations, some older applications that do not support Unicode file names might be unable to open files whose names contain non-ASCII characters after the files are synced, unshelved, or otherwise restored from a Unicode-mode server. |
See Troubleshoot user workstations in Unicode installations installations. |
|
You want to use separate workspaces (clients) and each of these needs to use a different character set |
Set a
different |
|
The files you check out need to be accessed by applications for which byte order is important. |
See Unicode character sets and Byte Order Markers (BOMs) or more information. |
|
You need to set |
See Controlling translation of server output for more information. |
|
The file is checked out using P4 Server client applications that handle Unicode environments in different ways. |
See Using other P4 Server client applications for more information. |
In each of these cases, you will need to explicitly set
P4CHARSET to an appropriate value or take some other action.
To get a list of the possible values for P4CHARSET, use the
command:
$ p4 help P4CHARSET
P4CHARSET while files are checked out. Submitting files using a different character set than the one used when syncing them can result in incorrect character translation.
Unicode character sets and Byte Order Markers (BOMs)
Byte order markers (BOMs) are used in Unicode files to specify the order in which multi-byte characters are stored and to identify the file content as Unicode. Not all extended-character file formats use BOMs.
To ensure that such files are translated correctly by the
P4 Server when the files are synced or submitted, you must set
P4CHARSET to the character set that corresponds to the
format used on your workstation by the applications that access them,
such as text editors or IDEs. Typically the formats are listed when you
save the file using the menu
option.
The following table lists valid settings for P4CHARSET for
specifying byte order properties of Unicode files.
| Client Unicode format | BOM? | Big or Little-Endian | Set P4CHARSET to | Remarks |
|---|---|---|---|---|
|
UTF-8 |
No |
(N/A) |
|
Suppresses P4 Server UTF-8 validation |
|
Yes |
|
|||
|
No |
|
|||
|
Yes |
|
|||
|
UTF-16 |
Yes |
Per client |
|
Synced with a BOM according to the client platform byte order |
|
Yes |
Little |
|
|
|
|
Yes |
Big |
|
||
|
No |
Per client |
|
||
|
No |
Little |
|
||
|
No |
Big |
|
||
|
UTF-32 |
Yes |
Per client |
|
Synced with a BOM according to the client platform byte order |
|
Yes |
Little |
|
||
|
Yes |
Big |
|
||
|
No |
Per client |
|
||
|
No |
Little |
|
||
|
No |
Big |
|
If you set P4CHARSET to a UTF-8 setting, the
P4 Server
does not translate text files when you sync or submit them.
P4 Server
does verify that such files contain valid UTF-8 data.
Controlling translation of server output
If you set P4CHARSET to any utf16 or
utf32 setting, you must set the
P4COMMANDCHARSET to a non-utf16 or
non-utf32 character set in which you want server output
displayed. "Server output" includes informational and error messages,
diff output, and information returned by reporting commands.
To specify P4COMMANDCHARSET on a per-command basis, use the
-Q flag. For example, to display all file names in the depot,
as translated using the winansi code page, issue the
following command:
C:\> p4 -Q winansi files //...
Using other P4 Server client applications
Different P4 Server client applications handle Unicode environments in different ways:
-
P4V prompts you to select a character set the first time you connect to a Unicode-mode server. P4V stores the selected character set with the connection and uses it for subsequent connections. You can also configure a global default character set, which is used automatically when connecting to Unicode-mode servers.
-
P4 for Eclipse prompts you to select a character set when connecting to a Unicode-mode server.
-
P4 Merge uses the character encoding configured through File > Character Encoding. When launched from P4V, P4 Merge uses the
P4CHARSETsetting configured in P4V instead of the character encoding specified in the P4 Merge preferences. -
P4 Plugin for Graphical Tools and P4EXP use the character set defined by the environment. These applications are not supported with Unicode-mode servers and will fail when connecting to them.