Pages

Monday, January 26, 2015

ESQL code to create mail with attachments using broker events

In this post I will provide an example on how to process events generated by a flow using the default IBM Integration Bus monitoring event capability.

The example will show how

  • to serialize a tree into BLOB using ESQL
  • how to send a mail with attachment
  • How to create/prepare the LocalEnvironment and Root for the EmailOutput node
  • The usage of the business event capabilities provided by IIB


The principle is simple:

  1. Configure a flow to generate an IIB event
  2. Create a subscription to this event with a WMQ Queue as endpoint
  3. Create a flow that consumes these events and send an email with attachments

The example in this post shows how to create mail with attachments using ESQL but this could be easily made using Java as well.

Configure a flow to generate an IIB event


The event generated as a well defined structure and the schema can be imported into a library using new model -> IBM predefined model.

Any nodes in a flow can be configured to generate events (Generating events in WebSphere Message Broker) that may contain context information and payload.

It is for example possible to configure a node to include the localEnvironment, ExceptionList and Message tree structure (under Root). These information will be placed into the IIB events under the folder "complexContent".
Note that the LocalEnvironment is reset when an exception occurs, so the data that would have been stored in this tree would be wiped when the message is propagated to the catch terminal of the input node (will be covered in a future post).

Finally it is also possible to include the full payload (as it was received) by selecting in the monitoring node properties "include payload as bitstream". The payload will then be included into the IIB events under "BistreamData".

Create a subscription

The IIB runtime is publishing the IIB events on the WMQ topic "$SYS/Broker/IBMIBus/Monitoring/#".
Using the WMQ Explorer you create a subscription to these events and select a destination queue:

 The flow that sends email

The flow is very simple: MQInput -> ComputeNode -> EmailOutput node
The compute node is used to create and configure the message that will be send using the emailoutput node. 
The node it self is configured the minimum properties: server:port, email to, from and security.
The rest will be provided by the code in ESQL (subject, body content and attachments).

In this example the complexContent included in the incoming business event is serialized into bitstream and will be send by mail as attachment.
The payload if present is also send as attachment.
The body of the mail is made of event origin data and using a DFDL to have a text document separated with CRLF.

The code is provided here after:




Friday, January 23, 2015

ESQL code Sample

In this post I will provide some example of ESQL code that could be useful.

The codes are made available using Gist.

TREE --> XML 

In the following gist, I provide an example on how to create a XML physical representation of a IIB in memory tree.
The principle is the following
  1. Create an element of type XMLNSC parser
  2. Copy or create a IIB tree
  3. Use the function ASBITSTREAM to serialize the tree in bitstream using the parser (here XML)


BLOB --> XML

In the following gist, I provide an example on how to create a XML tree from a BLOB.
In the example the BLOB is provided as test in a hexadecimal representation. The code parses it to an XML and append it in the current XML.


Friday, January 16, 2015

XPATH in IIB

In this post, I will provide some example of xpath expression allowing to perform complex transformation within a GDM (Graphical Data Mapper).

I will append this post with new example that I would found that could be of any usage.

For this first post, I will use a sample message having the following structure

<?xml version="1.0" encoding="UTF-8"?>
<Q1:INVOICE xmlns:Q1="http://www.acme.be/acme"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://www.acme.be/acmeInvoice.xsd ">
<CUSTOMERID>ID001</CUSTOMERID>
<NAME>X</NAME>
<INVOICE_ITEM>
<STOCK>OUT</STOCK>
<EXI>IMPORT</EXI>
<CONTAINER>abc</CONTAINER>
<ISOCODE>2210</ISOCODE>
<QUANTITY>1</QUANTITY>
</INVOICE_ITEM>
<INVOICE_ITEM>
<STOCK>OUT</STOCK>
<EXI>EMPTY</EXI>
<CONTAINER>abcd</CONTAINER>
<ISOCODE>2210</ISOCODE>
<QUANTITY>2</QUANTITY>
</INVOICE_ITEM>
<INVOICE_ITEM>
<STOCK>OUT</STOCK>
<EXI>EMPTY</EXI>
<CONTAINER>abcde</CONTAINER>
<ISOCODE>4532</ISOCODE>
<QUANTITY>4</QUANTITY>
</INVOICE_ITEM>
<INVOICE_ITEM>
<STOCK>OUT</STOCK>
<EXI>EMPTY</EXI>
<CONTAINER>abcdef</CONTAINER>
<ISOCODE>4532</ISOCODE>
<QUANTITY>2</QUANTITY>
</INVOICE_ITEM>
<INVOICE_ITEM>
<STOCK>IN</STOCK>
<EXI>IMPORT</EXI>
<CONTAINER>CONTAINER</CONTAINER>
<ISOCODE>ISOCODE</ISOCODE>
<QUANTITY>2</QUANTITY>
</INVOICE_ITEM>
</Q1:INVOICE>


Sum, Count

It is possible to compute the sum of elements under a repeating node in one expression.
For example if you would like to make the sum of the "Quantity" field of all the INVOICE_ITEM nodes, you could do this by performing the following map:

XPath expression is "fn:sum($INVOICE_ITEM/QUANTITY)"

Count can be used in the same way. Count will provide the number of elements.

Predicates

Predicates allows to select only a subset of nodes based on a criteria.
In the above example, it may be possible to compute the sum of QUANTITY for INVOICE_ITEM having its STOCK child element equal to "OUT".
This could be done by using predicates.
The above xpath expression would be changed with:

fn:sum($INVOICE_ITEM[STOCK='OUT']/QUANTITY)

The predicates here is "[STOCK='OUT']".

It is possible to use an input element within the predicates. For example, imagine that you have a list of Stock. And you would like to make the quantity sum only for the STOCK that are in the list.
This could be achieved by providing as input the stocklist to the xpath transform:
With the XPath expression
fn:sum($INVOICE_ITEM[STOCK=$STOCK1]/QUANTITY)

It is possible to have more than one "where" clause. For example is we would like the sum of  Quantity for STOCK='OUT' and EXI='IMPORT' then the following Xpath could be used:

fn:sum($INVOICE_ITEM[STOCK='OUT' and EXI='IMPORT']/QUANTITY)

be careful that the "and" is case sensitive.

Distinct Values

Last useful xpath expression for this post. Distinct values.
Distinct values could be used to retrieve only elements having distinct values.
For example let's say that the input message where STOCK can have the value "IN', "OUT'.
In the example above there are 4 Invoice Item with Stock equals to "OUT" and only 1 with "IN".
The XPATH expression used here is "fn:distinct-values($INVOICE_ITEM/STOCK)"
The input is the INVOICE_ITEM repeating node and the output is a repeating element STOCK under STOCKLIST.
The Stocklist will be populated by STOCK element having the value "IN" and "OUT":

<STOCKLIST>
        <STOCK>IN</STOCK>
        <STOCK>OUT</STOCK>
</STOCKLIST>

If you need to sort on multiple elements then the option that you may have is first create a concatenation of these fields in a previous map and then use the distinct values xpath expression.
So for example if you need to have the list of invoice items having distinct values for STOCK and EXI, then you may create a first map that concatenates STOCK and EXI using the xpath expression "fn:concat" and place this list in the localenvironment and then in the next map use the disctinct values xpath expression
.







HA Cache deployment with IBM Integration Bus

I have been involved in project where customer is looking to have a possibility to cache data  in a high available way within the integration layer.

In this post I will provide some points that have to be take into account when designing the IIB deployment architecture in order to have a high available cache.

Introduction

IBM Integration Bus provides an out of the box caching mechanism based on WebSphere eXtreme Scale.
WebSphere eXtreme Scale provides a scalable, in-memory data grid. The data grid dynamically caches, partitions, replicates, and manages data across multiple servers.

This cache can be used to store reference data that are regularly accessed or to hold a routing table.

The cache is not enabled by default and it is really easy to enable it: the default configuration is activated by setting a configuration parameters through the IBM Explorer administration tool.
To activate the cache across different Integration Node instances, a XML configuration file (templates are provided) has to be defined.
More information on the cache can be found here What's new in the Global Cache in IBM Integration Bus v9

Specialized skills in extremes scale is not necessary in order to use the cache. There is however two important cache components that may be good to know

  • Catalog servers: component that is embedded in an integration server and that controls placement of data and monitors the health of containers. You must have at least one catalog server in your global cache.
  • Container servers: component that is embedded in the integration server that holds a subset of the cache data. Between them, all container servers in the global cache host all of the cache data at least once. If more than one container exists, the default cache policy ensures that all data is replicated at least once. In this way, the global cache can cope with the loss of container servers without losing data.
More information about terminologies can be found here
Global cache terminologies

Principle

The catalog and container servers are embedded in Integration Servers.

To have a high available cache the following is required:

  • At least two catalog servers have to be online: without catalog server it is not possible to reach the data in memory
  • At least two container servers have to be online: this is necessary to replicate data on two different location

One more important point to know: it is not possible to configure an integration node to host a catalog server when it is configured as a multi-instance Integration node.

Possible deployment architecture

If the target is an active/active deployment, the following architecture is possible:
In this architecture, the catalog server is deployed in one Integration Server on both side. The other Integration servers are used to host containers. 
To improve the performance, the catalog server would be placed on a dedicated Integration Server (separated from the containers). This is not required though, a catalog server may reside on the same server as a container.
If the license doesn't permit to have multiple Integration Server per Integration Nodes (standard edition), then you could create a separate Integration Node on the same server to host the catalog server. 

If the target deployment consists of multi-instance queue managers, because for example the message residing on MQ has to be quickly recovered, the following architecture is possible:

Due to the fact that a multi-instance Integration Node can't host a catalog server (a configuration restriction), it is necessary to define an extra Integration Node to hold the catalog server (Integration Node - Catalog). This Integration Node doesn't need to be high available.
The multi-instance Integration Nodes are configured to host the container servers. Two active integration nodes are required to provide an high available cache (replication made on two different servers).

Additional information

The first time that the cache is configured using multiple catalog servers, the cache will become operational when at least two catalog servers are started. If the catalog servers are defined in two Integration Nodes, these two nodes would need to be started before been able to use the cache. Once the cache has been activated (this can be checked by looking at the logs or administration events) it is possible to lose (or stop) catalog servers (at least one catalog server should stay online) without impact to the cache access. This can be useful if maintenance on one instance has to be performed.

The default configuration is to have a maximum of 4 container servers per Integration Nodes. But more containers can be configured by configuring the Integration Server manually using the IBM Integration Explorer.

There are no limitation in term of container servers that can participate to the global cache. If you require more memory, you can just add a new container in the system. There is no need to restart the whole system !! Or you could also access an external extreme scale component like XC10 (Deciding between the embedded global cache and an external WebSphere eXtreme Scale grid).

Integration Server roles can be changed using the IBM Integration Explorer. It is possible to define a policy configuration file and assign it to the Integration Node using the IBM Integration Explorer. Once this has been done, you can start the Integration Node to take the configuration into account. If the global cache policy of the Integration Node is changed using the IBM Integration Explorer to "NONE" when the Integration Node is running, the current configuration will be held even if the Integration Node is restarted. It is therefore possible to change the Integration Server roles afterwards. More information on how to set the roles are provided here How to fix integration server roles in a Global Cache configuration in IBM Integration Bus and WebSphere Message Broker V8.

References

Information about global cache


Tuesday, October 28, 2014

IBM Integration Bus sizing and performance

It's not the first time that I am been in a situation where I have been asked to evaluate how many cores would be necessary to run a load on IBM Integration Bus (IIB).

In this post I will provide you a way to make this evaluation based on public IIB performance reports.

Required information  

Before starting your evaluation, you first need to identify a set of flows that corresponds to ESB patterns that you would like to achieve: transformation, aggregation, protocols, logging, ...

You would like as well to define on what operating system the run time will run.

For these flows you will need to define the following information
  • Kind of messages: XML, non XML
  • Average size of messages
  • Load peak expressed in transaction/sec. I usually use the peak since this is the throughput that you would like to sustain. 
  • Operating system where the IIB will run
Note that the CPU type has also it's importance. Normally you would have to add a corrector factor to compensate the fact that you have a more efficient processor.

Performance reports

Once this information is available, you can start your performance evaluation computation.
For this we will use performance report that is publicly available. For IBM Integration Bus, go to the link message throughput. (good link to bookmark is the performance topic). For older version, you can find performance reports at the support pack link.
On the message throughput link select your operating system.

These performance reports have been measure using a processor: IBM xSeries x3690 X5 with 2 x Deca-Core Intel(R) Xeon(R) E7-2860.
For those that would like to know the weight associated to this processor, please have a look at the PVU information page. You will find that this processor corresponds to a 70 PVU processor (2 sockets with 10 cores).

You will find different type of patterns that have been tested: aggregation, coordination request/reply, message routing, ...
For each of these tests there is a table providing the message rate, the CPU busy and the CPU ms/msg.

The information that we will keep are the 
  • message rate
  • the CPU ms/msg

Performance evaluation

How can we use these information?

Let's take an example!

Imagine that you would like to deploy a flow doing aggregation with an average non-persistent message size of 2 kb.
This flow will process 460000 messages a day but there is a peak! During one hour in the day the flow has to sustain 720000 messages.

We would like to size the bus such a way that he will sustain the peak. This corresponds to a throughput of 200 trs/sec.
Lets have a look to the tables available on the performance report page. There is one for the aggregation:

Non PersistentFull Persistent
Msg SizeMessage Rate% CPU BusyCPU ms/msgMessage Rate% CPU BusyCPU ms/msg
256b4020.199.34.941476.912.45.215
2kB3430.698.35.732707.819.75.559

For 2kb non persistent message, the processing of one message takes 5.732 ms of CPU.

The processing of 200 transactions per seconds would take 200 (msg/sec) *5.732 (CPU ms/msg) = 1146.4 (msg/sec)*(CPU sec/1000 * 1/msg) = 1.1464 CPU. 

So you would need to at least 2 processors to process this load. The total load of the machine with 2 CPU would be 57.32%.

It is not recommended to have a processor loaded at more than 70 % but here the 2 CPU will do the deal.

Of course you could have more complex situation such MQ, JMS, database, routing....

It is possible to approach more complex situation by adding the CPU load for each scenarios.
If my integration needs to be exposed as web service using SOAP messages, performs routing and transforms messages, I would approximate my CPU load by adding the corresponding load factor: 2.633 (SOAP) + 0.488 (routing) 0.896(transformation).

This is not ideal since each scenarios include the protocol processing like MQ for routing.
In the previous reports that are available in the support pack page, more tests have been done and it is possible to build a spreadsheets to isolate the processing of each mediations.
For example there is a test made for MQInput-MQOutput (x ms/msg), MQI-Transformation-MQO (y ms/msg) and JMSIn-JMSout (z ms/msg). For a JMSI-Transformation-JMSO integration flow, you could therefore approximate the CPU load using: y ms/msg - x ms/msg + z ms/msg.
The performance in IIB V9 is higher than for V8, so using the old performance reports would provide you a good estimation with a security factor.

References

message throughput information about IIB in function of the operating system.
performance information, tips and hints.
support pack old performance reports.
PVU information page to find the PVU for a defined CPU.

Tuesday, October 7, 2014

MQ Cluster demistified

I would like to provide through this post some highlighting on the MQ clustering: what is it, what does it brings and how to set it up.

MQ cluster is not about data replication or make data high available, it's about making object definitions known and available in a group of queue managers and providing a workload management capability.

 The "what is it"


Put it simply, a MQ Cluster is a collection of queue managers that have been defined to be part of a group called a cluster.
Within this group, queue managers know all objects that are shared within the cluster. Objects are like queues, topics.
Within a cluster all queue managers know how to send messages to the target clustered destination. The message transfer are handled automatically by the queue manager in the cluster.

The "what does it brings"

The cluster provides the following main advantages:
  • Simplify administration for distribute queueing. When a message has to send to a destination located to another queue manager, the queue has to be simply shared in the cluster on the remote queue manager that's it. The transfer of the message from the queue manager where the application has put the message to the remote queue manage is carried out by the cluster: no remote queues, no transmission queues, no channels. 
  • Workload balancing: when multiple instances of queues are defined, the cluster can balance the message distribution to all instances shared in the cluster. Priority and weight can be defined for specific queue managers. 
  • Improve availability and simplify maintenance: when multiple instance of queues are defined and one queue manager that hosts this queue is not available, the messages are automatically routed to the remaining queues. This increases the availability since messages can still be processed but it also simplifies the maintenance as the queue manager is known to be stopped in the cluster. 

Some disadvantages that I have in mind:
  • Queue manager in a cluster are directly connected with each others. It increases the number of channels which means consume more resources. 
  • Less control on how the messages are sent from one queue manager to another 
  • Less flexibility in the queue name resolution compared to the distributed queueing (alias, remote queues, ...). 

The "Principle" 

Let's first discuss about queue destinations.
When an application tries to put a message to a queue and this queue is not known by the local queue manager (the queue manager that the application is connected to), the queue manager queries the cluster (if it is part of one) to know if this queue is known.
If the queue is shared within the cluster, the information about its location is provided to the requestor queue manager. The local queue manager handles the message by putting it in the cluster transmission queue with a transmission header containing all the information to route it to the correct destination. The message is then sent by the local queue manager directly to the target queue manager.

Be aware that an application can get messages from local queue only, cluster defined or not. Local queues are queues defined on the queue manager that the application is connected to.
The story is a little bit different for topics. Indeed topic is used to define a destination where an application can publish or subscribe.
When a topic is shared within a cluster, it is known by all queue managers in this cluster.
Therefore if an application subscribes to a cluster topic it will receive all messages published to this topic even though the topic has been created in another queue manager.

The "how to make it" 

You will be puzzled how simple it is to create a cluster !!

Before going further there is one think to know about: the repositories. A repository is a store used by a queue manager to store cluster object definitions.

There are two types of repository: full and partial.
A full repository is a repository that contains cluster object definition from the whole cluster. When an object is shared in a cluster, the object definition is sent by this queue manager to the queue manager that holds the full repository.
A partial repository holds information about cluster object definitions that have been resolved from the queue manager that holds this repository. When an application puts a message for a remote queue shared in the cluster, the local queue manager requests the destination object information to the queue manager holding the full repository. This information is then stored in its local repository. Hence the name "partial repository" as it does not store all object definitions of the whole cluster.

In order to exchange these object definitions, queue managers use cluster channels:
  • a cluster receiver channel to receive object definitions 
  • a cluster sender channel to send object definitions to the queue manager(s) holding the full repository. 
Two rules about the repositories:
  • A queue manager can hold either a full repository or a partial repository (not both) 
  • Two (or one if not in production) full repositories is enough. All the other queue manager would hold a partial repository only. 
With this in mind, let's have a look how to make it !

The steps provided here after has to be performed in the same order.

For full repository
  1. Set the queue manager property "REPOS" to the name of the cluster. This will tell to the queue manager that he will host a full repository. A runmqsc command would be
    ALTER QMGR REPOS('myCluster')
     
  2. Create cluster receiver channel to define how to connect to this queue manager: 
    DEF CLUSRCVR(TO.MyQMgr) 
    CONNAME('ipaddressOfThisQMgr(portNumberOfThisQMgr)') 
    CLUSTER('myCluster')
  3. Optionally cluster sender channels can be created to point to another full repository ONLY. 
 For partial repository
  1. Create cluster receiver channel to define how to connect to this queue manager:
    DEF CLUSRCVR(TO.MyQMgr)
     CONNAME('ipaddressOfThisQMgr(portNumberOfThisQMgr)') 
    CLUSTER('myCluster')
  2. Create a cluster sender channel to define how to reach a queue manager holding a full repository (not partial).
    DEF CLUSSDR(TO.FullRepoQMgr) 
    CONNAME('ipaddressOfFullRepoQMgr(port)') 
    CLUSTER('myCluster')
That's it !!
Now you can create a queue and share it to the cluster "myCuster" that you just created.

References 

Information about the number of repositories in a cluster:
https://www.ibm.com/developerworks/community/blogs/messaging/entry/wmq_clusters_why_only_two_full_repositories?lang=en
Best practices (old presentation but the principle are still valid):
http://www.academia.edu/5513555/WMQCluster_Best_Practices
And this ten things for having an healthy cluster
https://www.ibm.com/developerworks/community/blogs/aimsupport/entry/ten_quick_tips_for_healthy_mq_cluster?lang=en
Very useful information about cluster questions
https://www.ibm.com/developerworks/community/blogs/messaging/tags/clustering?lang=en
Very good article that provides information about MQ HA
http://www.ibm.com/developerworks/websphere/library/techarticles/0505_hiscock/0505_hiscock.html

Wednesday, September 24, 2014

WebSphere MQ & IIB on Ubuntu

I found that more and more people like Linux and this is a great things.
IBM Integration Bus for developers (free of charge !!!) is available for Linux distribution.
That's cool... however I find out that the installation may be always straightforward especially on the last Linux 64 bit distribution.
You will find in this small post what are the steps that I followed to install the package.

First I would recommend to install each component separately:
  • WebSphere MQ
  • IBM Integration Bus
  • IIB Explorer

Download the IBM Integration Bus for developers package.

I have used the single package download.

WebSphere MQ


Installation

Installing software on Ubuntu usually entails using Synaptic or by using an apt-get command from the terminal. 
Unfortunately, there are still a number of packages out there that are only distributed in RPM format.
Ubuntu is Debian based and therefore uses.deb packages to install. If you want to install .rpm packages, you first should convert them into .deb packages with a conversion software such as alien. Then you can use gdebi or dpkg to install them.
Despite the large version number, alien is still (and will probably always be) rather experimental software. It has been used by many people for many years, but there are still many bugs and limitations

IBM distribute MQ as a set of RPMs and with what I just said, I recommend to use "rpm" and hence install the rpm package: ‘rpm’, ‘pax’ and ‘default-jre’ packages.
Detailed information can be found in the following technote:

In my installation I have made first 
1. Install the i386 libraries
sudo apt-get install libc6-i386
sudo apt-get install libgcc1:i386
sudo apt-get update 
2. Define the system configuration (provided here after)
3. Install the different rpm package as explaine at the previous link (for ubuntu 14.04 I had to use the --prefix /opt/mqm).
4. Define the WMQ configuration

System configuration

I would recommend to change the system configuration for Linux when using WebSphere MQ. I provide the information here after but detailed information can be found in the knowledge center ( Linux System requirements).

create /etc/sysctl.d/50-webspheremq.conf with the following:

kernel.shmmni = 4096
kernel.shmall = 2097152
kernel.shmmax = 268435456
kernel.sem = 500 256000 250 1024
fs.file-max = 524288
Make these changes live by running
sudo sysctl -p


Check that this has been correctly updated
cat /proc/sys/kernel/shmmni 

WMQ Configuration

You can use the setmqinst command to change the installation description of an installation, or to set or unset an installation as the primary installation.

Before executing the command, you would have to create the following directory “/usr/lib64”. 
Create this directory with the following command:
sudo mkdir /usr/lib64
This command sets the installation with an installation path of /opt/mqm as the primary installation:
sudo /opt/mqm/bin/setmqinst -i -p /opt/mqm

Add the user to the group mqm in order to be administrator:
sudo addgroup $USER mqm

IIBExplorer installation

Before doing any installation, you should first install the prereqs librairies (see corresponding chapter).

Install IBExplorer. Go to the folder “/extractedfolder/messagebroker_ia_developer/IBExplorer”.
If the installer is not executable, you would have to change using the following command:
chmod +x install.bin

launch the installation:
sudo ./install.bin

And start the IBM Integration Toolkit with the following command:
sudo strmqcfg -clean

IBM Integration Toolkit

Before doing any installation, you should first install the prereqs librairies (see corresponding chapter).

The installation is then straightforward, just launch the installToolkit.

sudo ./installToolkit.sh

IBM Integration Bus

The runtime is installed in a very straightforward way.
Just go in the directory where the tar has been unzipped and launch the setup.

Intalling Eclipse prereqs

The installation manager, the IBM Integration Toolkit and IBM Integration Exporer requires to have 32bit libraries. I have found issues to run the installation without installing the "ia32-libs".
In order to install this library I have used the following command:
 
sudo -i
cd /etc/apt/sources.list.d
echo "deb http://old-releases.ubuntu.com/ubuntu/ raring main restricted universe multiverse" >ia32-libs-raring.list
apt-get update
apt-get install ia32-libs
rm /etc/apt/sources.list.d/ia32-libs-raring.list
apt-get update
exit
sudo apt-get install gcc-multilib

And the following libs:
apt-get install libc6-i386
apt-get install libgcc1:i386