LinkedIn

Tuesday, January 27, 2009

Optimizing XQueries

Well can we imagine programming in the SOA world without knowledge of XML Technologies. As a matter of fact if we are working on ALSB and ALDSP then the knowledge of XPaths and XQueries is of prime importance. Here i would be discussing the optimal practice of writing XQueries.

Explanations of each of the following tips can be found at the end of the article.

DON'TS
Here are a few things that we must seek to avoid:
  1. Don't use eval ()
  2. Don't evaluate expressions several times over, and avoid redundant expressions.
  3. Don't use //
  4. Don't query constructed document fragments
    DOS
    Here are some recommendations for optimization:
    1. Minimize the execution of queries based on a given search expression. Try instead to use navigation paths based on the parent, children and siblings of a node which has already been retrieved
    2. Make appropriate use of indexes adapted to your search criteria.
    3. Code Quality
      TODO
      1. Put $Id$ inside a comment at the top internal documentation of HTTP parameters
      2. Document in Xquery the argument types and the return type
      3. Use meaningful names for variables and functions, without abbreviations, and avoid ambiguous terms
      4. Use Javadoc-style tags as in XQDOC ( http://www.xqdoc.org/ ) : @param, @return
      5. Keep data retrieval separate from result construction
        _______________________________________________________________________________________________

        EXPLANATIONS

        Don't use eval ()

        The snag is, the arguments to the eval () function can't be cached. Beyond that, using eval () leads to a style of programming that's hard to read and to debug. And eval () can always be replaced by a standard expression.

        Don't evaluate expressions several times over and avoid redundant expressions

        Xquery doesn't perform any analysis or optimization of queries akin to what a Java compiler does. So no refactoring of repeatedly-evaluated expressions, no elimination of code that won't be executed, etc. Pay particular attention to repeatedly evaluated expressions, they should be evaluated once only and the result placed into a variable, which also makes for more readable code.

        Don't use //

        $a//b causes a complete traversal of all nodes of which $a is the root in search of an element b. In most cases the location of b is fairly precisely known, and so would be better to specify it.

        Don't query constructed document fragments

        A typical example (to avoid):
        let $e := content (: $e is a constructed document fragment :)
        let result := $e/b/text()

        Minimize the execution of queries based on a given search expression.

        A query like

        res := collection("/db/projects") /a/b [ id = $val ]

        causes a complete scan of an entire collection. Admittedly, queries like this are at the heart of an XQuery (and account for most of its execution time). But once the result $res has been retrieved, it can be efficiently used as a starting point for navigation to its parent, siblings and children:

        $a: = $res / parent::a
        $next-sibling: = $a / next-sibling:a

        Make appropriate use of indexes adapted to your search criteria.

        There are currently three types of user-configurable indexes in Xquery. All require pre-indexation either of the base collection or of specified node-sets in sub-collections.
        • The fulltext index, which indexes lexical tokens ("words" in Western scripts). Indexation can be configured to include or exclude nodes specified using a limited subset of XPath
        • Typed indexes over nodes specified by a limited subset of XPath (called "range indexes" because they permit queries referring to a range of numerical values)
        • Indexes by tag name ("Qname index") http://wiki.exist-db.org/space/jmvanel/New+index+by+QName
        Index 2. is slower than 3., but has two advantages

        The request code doesn't have to be changed in order to use the index with 2, there is no danger of getting wrong results if the indexation hasn't been done.

        Index 3 lacks these advantages, but is almost as fast as a relational database. Such an index cannot be constrained by an XPath, but only by a tag name. Both index and and index 2 are typed (integers or strings), and allow matching by criteria of equality or inequality (comparison).

        Document in XQuery the argument and return types

        Don't write :

        declare function local:add($n, $m) {
        $n + $m
        };
        This is more explicit and auto-documenting. And for the same price you get run-time arguments checking. If you know for sure the types you manipulate, declare them !

        declare function local:add($n as xs:integer, $m as xs:integer)
        as element(result) {
        $n + $m
        };

        Also Keep data retrieval separate from result construction. It is good to create variables of child element nodes that are used in result construction rather than retrieving them every time from the root.

        Monday, January 12, 2009

        Open Source SOA Implementation

        Hardly would there be anybody who hasnt heard or been a part of the SOA hype over the couple of years. It seems as if an entire industry is converging towards this amazing concept of sharing, interoperability and plug-ability. For long software architects for medium and large businesses have always had to go through a nightmare while designing systems that were heterogeneous. Well this is not a blog to present the advantages and disadvantages of SOA, its various pros and cons and the ideas revolving around its concepts.

        Neverthless there has been a host of vendors that have leaped to embrace SOA and have created an entire product suite/stack for implementing it. For instance there is BEA with its entire Weblogic and Aqualogic suite, IBM with its Websphere Development Studio, Tibco, WebMethods and a host of other proprietary vendors.

        I have however tried to make a presentation on how to implement using open source tools and technologies. Here is a slide that evaluates the various options available for an open source implementation of SOA.

        Find it here.

        Monday, December 15, 2008

        Creating Weblogic 9.2 Cluster (With Remote Managed Servers) with WLST

        Overview


        This section describes how to setup a WebLogic 9.2 cluster on Linux servers using the BEA WebLogic Scripting Tool (WLST). WLST is a scripting tool installed with WebLogic that allows for command-line configuration of the WebLogic server. The set of scripts provided here along with these instructions will allow you to setup a WebLogic cluster in your environment with ease.

        Clustering Scripts Setup


        In preparation for running the WLST scripts to create a cluster, you will need to download the WLST Clustering Scripts (ZIP, 7KB) and unzip it to a directory. After that, configure the environment variables in all of the scripts to match your environment. A description of the variables that will need to be set can be found in the section Environment Variables to Configure in the WLST Scripts. After that, copy the files to the servers that you will be using in the cluster. You will then need to navigate to this directory on a server to execute any of the scripts on that server.

        Steps to Create the WebLogic Domain and Setup the Cluster


        On the Server hosting the AdminServer:

        Create a domain with AdminServer and a cluster of two managed servers

        WL_HOME/common/bin/wlst.sh createcluster.py DOMAIN_NAME CLUSTER_NAME
        Start the node manager and AdminServer WL_HOME/common/bin/wlst.sh startadminserver.py DOMAIN_NAME
        Setup JDBC Data Source WL_HOME/common/bin/wlst.sh createjdbc.py t3://AdminServerIP:AdminServerHttpPort CLUSTER_NAME
        Pack the domain WL_HOME/common/bin/pack.sh -managed=true -domain=DOMAIN_PATH -template=DOMAIN_TEMPLATE -template_name=DOMAIN_TEMPLATE_NAME

        On all servers not hosting the AdminServer:

        Unpack the domain Copy DOMAIN_TEMPLATE from the server hosting the AdminServer to the same directory on all other servers that are apart of the cluster and then execute the following command on the servers to create the base domain: WL_HOME/common/bin/unpack.sh -domain=DOMAIN_PATH -template=DOMAIN_TEMPLATE
        Enroll the node manager with AdminServer and then start it Execute the following command on the servers to setup the node managers and start them: WL_HOME/common/bin/wlst.sh enrollnodemanager.py t3://AdminServerIP:AdminServerHttpPort DOMAIN_NAME

        On the Server hosting the AdminServer:

        Deploy the Elastic Path code to the cluster WL_HOME/common/bin/wlst.sh deploy.py t3://AdminServerIP:AdminServerHttpPort DEPLOYMENT_NAME APPLICATION_PATH CLUSTER_NAME
        Start the cluster WL_HOME/common/bin/wlst.sh startcluster.py t3://AdminServerIP:AdminServerHttpPort CLUSTER_NAME

        Additional Useful Scripts


        To create another server in the cluster, execute this command after step 1 WL_HOME/common/bin/wlst.sh createmanagedserver.py DOMAIN_NAME CLUSTER_NAME
        To shut down server ServerName WL_HOME/common/bin/wlst.sh stopserver.py DOMAIN_NAME ServerName
        To remove a deployment WL_HOME/common/bin/wlst.sh undeploy.py t3://AdminServerIP:AdminServerHttpPort DEPLOYMENT_NAME

        Descriptions of Constants Used in Setup Steps


        BEA_HOME The BEA Home directory where files common to all BEA products are stored (eg. /opt/bea/)
        WL_HOME The WebLogic Server product installation directory (eg. /opt/bea/weblogic92/)
        AdminServerIP The IP address of the server that the AdminServer is on
        AdminServerHttpPort The port for the AdminServer to listen to http requests on
        DOMAIN_NAME The name of the clustered domain to be created (eg. epclusterdomain)
        DOMAIN_PATH BEA_HOME/user_projects/domains/DOMAIN_NAME (the path to the clustered domain)
        DOMAIN_TEMPLATE The filename of the domain template to be created or accessed (eg. BEA_HOME/user_templates/epclusterdomain_managed.jar)
        DOMAIN_TEMPLATE_NAME The descriptive name of the domain template to be created (eg. "EP Clustered Domain")
        CLUSTER_NAME The name of the cluster to create (eg. wlsCluster)
        DEPLOYMENT_NAME The name for the deployment of Elastic Path application code (eg. epsf_cluster_deployment)
        APPLICATION_PATH The path to where the Elastic Path application has been setup (eg. /home/build/ep_weblogic/com.elasticpath.sf/)

        Environment Variables to Configure in the WLST Scripts


        BEA_HOME The BEA Home directory where files common to all BEA products are stored (eg. /opt/bea/)
        WL_HOME The WebLogic Server product installation directory (eg. /opt/bea/weblogic92/)
        JAVA_HOME The root directory of the Java JDK install that is used to run WebLogic (eg. /opt/j2sdk)
        AdminServerIP The IP address of the server that the AdminServer is on
        AdminServerHttpPort The port for the AdminServer to listen to http requests on
        AdminServerHttpsPort The port for the AdminServer to listen to https requests on
        AdminServerPassword The password used to connect to the AdminServer as default user weblogic
        Machine1IP The IP address of the server hosting the first managed server in the cluster
        Machine1Name The machine name of the server hosting the first managed server in the cluster
        Server1HttpPort The port for the first managed server to listen to http requests on
        Server1HttpsPort The port for the first managed server to listen to https requests on
        Server1Name The name to use for the first managed server in the cluster (eg. epServer1)
        Machine2IP The IP address of the server hosting the second managed server in the cluster
        Machine2Name The machine name of the server hosting the second managed server in the cluster
        Server2HttpPort The port for the second managed server to listen to http requests on
        Server2HttpsPort The port for the second managed server to listen to https requests on
        Server2Name The name to use for the second managed server in the cluster (eg. epServer2)
        JdbcName The descriptive name of the JDBC data source (eg. EP)
        JndiName The JNDI name of the JDBC data source (eg. jdbc/epjndi)
        Url The JDBC connection URL (eg. jdbc:oracle:thin:@11.11.1.111:1111:ep)
        JdbcDriverName The name of the JDBC driver (eg. oracle.jdbc.OracleDriver)
        DbUserName The username for accessing the database
        DbUserPassword The password for accessing the databas

        Thursday, September 25, 2008

        Stripping XML Namespaces using XQuery

        A lot of the times I had found people complaining about unneccessary namespaces that flood their xml. Offcourse namespaces are very important but during certain times when we have to process complex and large xmls using XPath it becomes necessary to strip the namespaces. Here is a demonstration of a small XQuery code that is used to strip any samespaces from an XML input.

        (:

        =====================================
        Description: xquery to remove namespaces from an xml
        @author Arun Pareek
        =====================================
        $Resource/XQUERY/stripNamespace.xq  $

        :)

        declare variable $inputRequest as element() external;
        declare function strip-namespace($inputRequest  as element()) as element()
        {
        element {xs:QName(local-name($inputRequest ))}
        {
        for $child in $inputRequest /(@*,node())
        return
        if ($child instance of element())
        then strip-namespace($child)
        else $child
        }
        };

        document
        { strip-namespace($inputRequest )}

        Before explaining the code let me give a small example of how this piece of code works.

        Input XML with Namespaces

        <TestXML xmlns:sp-en="sch.soap.org">
        <ZBAPI_COIB.ZBUR xmlns:sp-en="sch.soap.org">
        <I_OUTPUT xmlns:sp-en="sch.soap.org"/>
        <RETURN xmlns:sp-en="sch.soap.org">
        <item xmlns:sp-en="sch.soap.org">
        <TYPE xmlns:sp-en="sch.soap.org">S</TYPE>
        <CODE xmlns:sp-en="sch.soap.org"/>
        <MESSAGE xmlns:sp-en="sch.soap.org"/>
        <LOG_NO xmlns:sp-en="sch.soap.org"/>
        <MESSAGE_V1 xmlns:sp-en="sch.soap.org"/>
        </item>
        <item xmlns:sp-en="sch.soap.org">
        <TYPE xmlns:sp-en="sch.soap.org">S</TYPE>
        <CODE xmlns:sp-en="sch.soap.org"/>
        <MESSAGE xmlns:sp-en="sch.soap.org"/>
        <LOG_NO xmlns:sp-en="sch.soap.org"/>
        <LOG_MSG_NO xmlns:sp-en="sch.soap.org">000000</LOG_MSG_NO>
        <MESSAGE_V1 xmlns:sp-en="sch.soap.org"/>
        </item>
        </RETURN>
        </ZBAPI_COIB.ZBUR>
        </TestXML>

        And the Output of the Xquery is

        <TestXML>
        <ZBAPI_COIB.ZBUR>
        <I_OUTPUT/>
        <RETURN>
        <item>
        <TYPE>S</TYPE>
        <CODE/>
        <MESSAGE/>
        <LOG_NO/>
        <MESSAGE_V1/>
        </item>
        <item>
        <TYPE>S</TYPE>
        <CODE/>
        <MESSAGE/>
        <LOG_NO/>
        <LOG_MSG_NO>000000</LOG_MSG_NO>
        <MESSAGE_V1/>
        </item>
        </RETURN>
        </ZBAPI_COIB.ZBUR>
        </TestXML>

        As you can see the code is very much self-explanatory. The Xquery snippet expects an xml input to it. The Xpath entity crawls through each of the elements in the input and returns just the child elemnt and discards the namespace. A very much similar to recursive fucntions in C/C++.

        Enjoi :-)