Sunday, 17 May 2015

Repository Cached Items are no more puzzle


While analyzing cache behaviour for our application. I came to know one more interesting thing related to repository cache. That is to print the content of item-cache.



The <dump-caches> tag can be used to print out the contents of the item cache for one or more item descriptors.

There are two attributes in this tag.
  
   1. dump-type It takes  one of the values from below set.
         - debug : To display  cache items.
         - queries : Creates a log entry consisting of the <load-items> tag that is used to reload the cache.
         - both : Combines the output of debug and queries.
   2. item-descriptors  A comma-separated list of one or more item descriptor names. If no item descriptors are specified, all item descriptor caches are exported.
 

Wednesday, 13 May 2015

ATG Repository Item Cache Interesting


While analyzing repository item cache related issues. I found interesting point about item cache size configuration in the case of item descriptor inheritance. 

Link :
http://docs.oracle.com/cd/E26180_01/Platform.94/RepositoryGuide/html/s1010cacheconfiguration01.html

How item cache size works : It is the maximum number of items of this type to store in the item cache. When this maximum is exceeded, the oldest items are removed from the cache.

Here is the fact : Within an inheritance tree, the last-read item descriptor’s item-cache-size setting applies to all item descriptors within the inheritance tree. In order to ensure the desired item cache size, be sure to assign the same item-cache-size to all related item descriptors 

Default Item Cache Value : 1000

Monday, 20 April 2015

ATG Calendar Scheduler interesting

Today I am going to explain very interesting point about ATG schedulers.



To configure the scheduler in ATG we need to assign value to schedule property of the scheduler.
 There are 3 types of formats for this property.
  1. Relative Schedule.
  2. Periodic Schedule.
  3. Calendar Schedule.
Out of above mentioned formats Calendar Schedule is most confusing.Here is the syntax for Calendar Schedule.

Syntax : calendar <months> <dates> <days of week>  <occurrences in month> <hours> <minutes>

Example schedule 

calendar * * 1 * 22 0

Here the * in date field will invoke this scheduler every day. 

Thumb rule is

A * entry selects all values for that field. A period (.) selects no values.

calendar * . 1 * 22 0

Above scheduler will run weekly.
 

Saturday, 11 April 2015

Index xml File into Endeca

Here is the small  example with steps to index xml file into Endeca.

1. First of all configure xml file in Record Adapter of pipeline.

2. Create your xml file.Below is the sample xml file content. Property name is the name of the source property defined in the property mapper. This is Endeca Record XML format. XML adapters consumes data in this format without transformation, other xml formats cannot be read by the data foundry.To support these situations, an XSLT transformation can be applied to the source data to convert it into Endeca Records XML, which the Data Foundry can read.In this case configure xslt in transformer tab of record adapter configuration.

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE RECORDS SYSTEM "records.dtd">
<RECORDS>
     <RECORD>
        <PROP NAME="SampleProperty1">
           <PVAL>12</PVAL>
        </PROP>
        <PROP NAME="SampleProperty2">
           <PVAL>Jagdev</PVAL>
        </PROP>
        <PROP NAME="SampleDimension1">
           <PVAL>1029</PVAL>
        </PROP>
     </RECORD>
</RECORDS>


3. Run baseline indexing and see the resutls in jspref.


Tuesday, 31 March 2015

Create And Provision Endeca Application with the Deployment Template


Here are the steps to create an Endeca application from scratch and provision it using deployment template.

Application Creation 
 
      1. Open command line and change directory to ..\...\ToolsAndFrameworks\3.1.2  
          \deployment_template\bin and invoke deploy script. 

           For ex. C:\Endeca\ToolsAndFrameworks\3.1.2\deployment_template\bin\deply.bat
 
      2. Provide configuration parameter/values as script prompts you. First of all it will confirm
          IAP version. Then it will prompt following information.
  • Application Name.
  • Application Deployment directory.
  • EAC port.
  • Workbench port.
  • Live dgraph port.
  • Authoring Dgraph port.
  • Log Server port.
Screen  Captures for application creation.

For port values you can use any available port on your machine, or use default one.At the end you will get message that appplication deployed successfully.

Now application creation is done. It is time to provision it. Once application provisioned it will be available for configuration in Endeca Workbench.

To provision an application.

Go to \control directory of your application you created earlier (above steps).
Invoke initialize_services.bat or initialize_services.sh script. This script won't prompt for any values.

Here is the screen capture of application provisioning.
 
 

Sunday, 15 March 2015

Speed Up ATG Application Using Caching

Here I am going to introduce repository level caching mechanism in ATG. Caching is very critical to application performance.


The thumb rule for caching is 

You should design an application so it requires minimal access to the database and ensures data integrity. 

Each item descriptor in SQL repository has its own 2 types of cache.
  1. Item Cache
  2. Query Cache 
 1. Item Cache : The item cache holds property values for repository items. It is indexed by the repository item IDs.

2.  Query Cache : The query cache holds the repository IDs of items that match particular queries in the cache.

 Advantages of having item and query cache at each item descriptor level.
  • Set caching size for each item type separately.
  • Flush cache for each item type separately (selective cache invalidation). 
 
How is work ?

Repository queries are performed in two passes, using two separate SELECT statements.

1. Repository id fetching (query cache) : The first statement gathers the IDs of the repository items that match that query. Repository first examines the query cache whether same query is already cached or not. In the case query is already cached, it returns matching ids from query cache. In this case no query will be fired on Database. If this query is not cached, then repository fires query on database. Create entry in query cache for this query and return the results. It is very useful when repeated queries are common.

2. Repository Item fetching (Item cache) : The SQL repository then examines the result set from the first SELECT statement and finds any items that already exist in the item cache. A second SELECT statement retrieves from the database any items that are not in the item cache. 

Interesting fact about query cache : When application executes query using id parameter.
For example id="1234"
In this case ATG is not going to create entry in query cache. As query parameter and id returned from query is same. It bypass the first query, and directly work on second query and item cache.

Cache Tuning :

1. Query Cache Tuning : It is generally safe to set the size of the query cache to 1000 or higher. Query caches only contain the query parameters and string IDs of the result set items, so large query cache sizes can usually be handled comfortably without running out of memory.A query whose parameters are subject to less frequent changes is a good candidate for caching.

2. Item Cache Tuning : Item cache size should be large enough to accommodate the number of items in the repository. For example repository contains 50000 SKUs. Then sku item cache size should be 50000. Item cache size need set properly as it requires more space(memory) as compare to query cache. You need to test application performance with different item cache size based on hardware availability.

Good Luck..!!

Saturday, 7 March 2015

Resolving Performance Puzzle

It is challenging for beginners to start performance tuning of an application. I was in same situation when got my first performance tuning assignment. Before starting with performance tuning we need to understand the aspects of it. 


In simple words Performance tuning is the improvement of system performance. Most systems will respond to increased load with some degree of decreasing performance. The core of performance tuning is the performance testing.


Here is the list of steps to execute systematic tuning.
  1. Assess the issue, get the numbers (to baseline the performance) to measure the performance (Mostly available/defined in Non Functional Requirement [NFR] document).
  2.  Measure the performance of system.
  3. Identify the bottlenecks (Part of the system that is critical for performance).
  4. Modify the system update/remove the bottleneck.
  5. Measure the system performance after above modification.
  6. In the case modification improves the performance adopt it,otherwise revert the above modification.
  7. Repeat from steps 2 to 6 in cycles to improve the performance requirement.
These are high level steps. The most confusing/challenging question comes into mind is. 

Where (which part of application) to start performance testing ?.

Most of the web applications are modularized/divided into 3 major parts.
  1. Database.
  2. Back-end code (Including third party calls). 
  3. Front-end.


Its always better to start from database and end at front-end.