A Developer's Diary

Showing posts with label Encoding. Show all posts
Showing posts with label Encoding. Show all posts

Nov 4, 2012

Removing BOM character from a string in Java

The below code checks if the BOM character is present in a string. If present, removes and prints the string with skipped bom.

import java.io.UnsupportedEncodingException;

/**
 * File: BOM.java
 * 
 * check if the bom character is present in the given string print the string
 * after skipping the utf-8 bom characters print the string as utf-8 string on a
 * utf-8 console
 */

public class BOM
{
    private final static String BOM_STRING = "Hello World";
    private final static String ISO_ENCODING = "ISO-8859-1";
    private final static String UTF8_ENCODING = "UTF-8";
    private final static int UTF8_BOM_LENGTH = 3;

    public static void main(String[] args) throws UnsupportedEncodingException {
        final byte[] bytes = BOM_STRING.getBytes(ISO_ENCODING);
        if (isUTF8(bytes)) {
            printSkippedBomString(bytes);
            printUTF8String(bytes);
        }
    }

    private static void printSkippedBomString(final byte[] bytes) throws UnsupportedEncodingException {
        int length = bytes.length - UTF8_BOM_LENGTH;
        byte[] barray = new byte[length];
        System.arraycopy(bytes, UTF8_BOM_LENGTH, barray, 0, barray.length);
        System.out.println(new String(barray, ISO_ENCODING));
    }

    private static void printUTF8String(final byte[] bytes) throws UnsupportedEncodingException {
        System.out.println(new String(bytes, UTF8_ENCODING));
    }

    private static boolean isUTF8(byte[] bytes) {
        if ((bytes[0] & 0xFF) == 0xEF && 
            (bytes[1] & 0xFF) == 0xBB && 
            (bytes[2] & 0xFF) == 0xBF) {
            return true;
        }
        return false;
    }
}


The following stackoverflow article talks in detail about detecting and skipping the BOM

Read more ...

Byte Order Mark (BOM) character

BOM character
1. BOM is a Unicode character used to identify the endianness of the text file or stream.
2. The UTF-8 representation of the BOM is the byte sequence 0xEF, 0xBB, 0xBF.
3. A text editor using ISO-8859-1 as character encoding will display the characters  for BOM.
4. BOM has no meaning in UTF-8 apart from signalling that the byte stream that follows is encoded in UTF-8

import java.nio.charset.Charset;

/**
 * File: BOM.java
 * 
 * The following class converts a string having bom character
 * from ISO-8859-1 encoding type to UTF-8 and back
 */
public class BOM
{
    public static void main(String[] args) throws Exception
    {
        System.out.println("Default Encoding: " + Charset.defaultCharset());

        //
        // Displays a simple string with bom prepended.
        // Uses system default character encoding
        //
        String bomString = "Hello World";
        System.out.println(bomString + " Length: " + bomString.length());

        //
        // convert string with bom character to utf string
        //
        byte[] byteArrayISO = bomString.getBytes("ISO-8859-1");
        String utfString = new String(byteArrayISO, "UTF-8");
        System.out.println(utfString + " Length: " + utfString.length());

        //
        // convert the utf string back to windows character encoding
        //
        byte[] byteArrayUTF = utfString.getBytes("UTF-8");
        String winString = new String(byteArrayUTF, "ISO-8859-1");
        System.out.println(winString + " Length: " + winString.length());
    }
}

Output of the above program when run on a UTF-8 supported console
$ java BOM
Default Encoding: windows-1252
Hello World Length: 17
Hello World Length: 14
Hello World Length: 17

Read more ...