Few days ago, I got a request to explain how to generate the SCHEMA file of a table ? So here we goes.....
Something about DataStage, DataStage Administration, Job Designing,Developing, DataStage troubleshooting, DataStage Installation & Configuration, ETL, DataWareHousing, DB2, Teradata, Oracle and Scripting.
Showing posts with label Metadata. Show all posts
Showing posts with label Metadata. Show all posts
Tuesday, October 14, 2014
Auto Generate Table Schema in DataStage
Few days ago, I got a request to explain how to generate the SCHEMA file of a table ? So here we goes.....
Monday, July 21, 2014
Navigating the many paths of metadata for DataStage 8
Source : Navigating the many paths of metadata for DataStage 8
Looking at the methods for importing metadata table definitions into DataStage 8 ETL jobs.
All of the metadata import methods of DataStage 7 are in DataStage 8 and all execute in the same way. Developers familiar with previous versions will be right at home! What is tricky to come to terms with are all the new ways to get metadata into DataStage 8 and the Metadata Server. The Metadata Server can provide reporting on metadata from products outside of DataStage such as BI tools so in some cases you might be importing the metadata for reporting and not for ETL.
The list that follows covers just the techniques for importing metadata to be used by DataStage jobs.
Looking at the methods for importing metadata table definitions into DataStage 8 ETL jobs.
All of the metadata import methods of DataStage 7 are in DataStage 8 and all execute in the same way. Developers familiar with previous versions will be right at home! What is tricky to come to terms with are all the new ways to get metadata into DataStage 8 and the Metadata Server. The Metadata Server can provide reporting on metadata from products outside of DataStage such as BI tools so in some cases you might be importing the metadata for reporting and not for ETL.
The list that follows covers just the techniques for importing metadata to be used by DataStage jobs.
Tuesday, June 17, 2014
FastTrack Makes Your DataStage Development Faster
IBM introduced a tool called FastTrack that is a source to target mapping tool that is plugged straight into the Information Server and runs inside a browser.
The tool was introduced with the Information Server and is available in the 8.1 version.
As the name suggests IBM are using it to help in the analysis and design stage of a data integration project to do the source to target mapping and the definition of the transform rules. Since it is an Information Server product it runs against the Metadata Server and can share metadata with the other products and it can run inside a browser.
I have talked about it previously in New Product: IBM FastTrack for Source To Target Mapping and FastTrack Excel out of your DataStage project but now I have had the chance to see it in action on a Data Warehouse project. We have been using the tool for a few weeks now and we are impressed. It’s been easier to learn than other Information Server products and it manages to fit most of what you need inside frames on a single browse screen. Very few bugs and it has been in the hands of someone who doesn’t know a lot about DataStage and they have been able to complete mappings and generate DataStage jobs.
I hope to get some screenshots up in the weeks to come but here are some observations in how we have saved time with FastTrack:
Monday, February 17, 2014
Datastage Coding Checklist
- Ensure that the null handling properties are taken care for all the nullable fields. Do not set the null field value to some value which may be present in the source.
- Ensure that all the character fields are trimmed before any processing. Normally extra spaces in the data may lead to some errors like lookup mismatch which are hard to detect.
- Always save the metadata (for source, target or lookup definitions) in the repository to ensure re usability and consistency.
Wednesday, February 12, 2014
Interview Questions : DataWareHouse - Part 4
What are the types of Synonyms?
There are two types of Synonyms Private and Public
What is a Redo Log?
The set of Redo Log files YSDATE, UID, USER or USERENV SQL functions, or the pseudo columns LEVEL or ROWNUM.
What is an Index Segment?
Each Index has an Index segment that stores all of its data.
Explain the relationship among Database, Table space and Data file?
Each databases logically divided into one or more table spaces one or more data files are explicitly created for each table space.
Thursday, December 12, 2013
Interview Questions : DataStage - self-3
100 If 1st and 8th record is duplicate then which will be skipped? Can you configure it?
101 How do you import and export datastage jobs? What is the file extension? (See each component while importing and exporting).
102 How do you rate yourself in DataStage?
103 Explain DataStage Architecture?
104 What is repository? What are the repository items?
105 What is difference between routine and transform?
106 When you write the routines?
Thursday, November 14, 2013
DataStage Server Hang Issues & Resolution
Server hang issue can occurred when
1) Metadata repository database detects a deadlock condition and choose failing job as the victim of the deadlock.
2) Log maintenance is ignored.
3) Temp folders are not maintained periodically.
I will try to explain above three points in detail below:
1) Metadata repository database detects a deadlock condition and choose failing job as the victim of the deadlock.
2) Log maintenance is ignored.
3) Temp folders are not maintained periodically.
I will try to explain above three points in detail below:
Tuesday, November 12, 2013
ETL Job Design Standards - 1
When using an off-the-shelf ETL tool, principles for
software development do not change: we want our code to be reusable, robust,
flexible, and manageable. To assist in the development, a set of best practices
should be created for the implementation to follow. Failure to implement these
practices usually result in problems further down the track, such as a higher
cost of future development, increased time spent on administration tasks, and
problems with reliability.
Although these standards are listed as taking place in ETL
Physical Design, it is ideal that they be done before the prototype if
possible. Once they are established once, they should be able to be re-used for
future increments and only need to be reviewed.
Listed below are some standard best practice categories that
should be identified on a typical project.
Labels:
database
,
DataStage
,
design
,
environment
,
Errors
,
ETL
,
handling
,
Job
,
Link
,
managers
,
Metadata
,
names
,
notification
,
Optimizing
,
parameter
,
process
,
reusability
,
stages
Friday, September 20, 2013
Failure to connect to DataStage services tier: invalid port
When you attempt to start one of the DataStage clients, the following message is displayed:
Failed to authenticate the current user against the selected Domain: Could not connect to server [servername] on port [portnumber].
Wednesday, March 13, 2013
All about 000 - 421 : DataStage Certification Exam Test Preparation
1.
DataStage v8 Configuration (5%)
- Describe
how to properly configure DataStage V.8.0.
- This
is kind of vague but focus on how DataStage 8 gets attached to a Metadata
Server via the Metadata Console and how security rights are set up.
- Read
up on configuring DB2 and Oracle client and ODBC.
- Get to know the dsenv file. Read the DataStage Installation Guide for post-installation steps.
Labels:
421
,
certification
,
Configuration
,
Data
,
database
,
DataStage
,
design
,
exam
,
Job
,
Metadata
,
Parallel
,
storage
,
transformation
Monday, February 04, 2013
14 Good design tips in Datastage
1) When you need to run the same sequence of jobs again and again, better create a sequencer with all the jobs that you need to run. Running this sequencer will run all the jobs. You can provide the sequence as per your requirement.
2) If you are using a copy or a filter stage either immediately after or immediately before a transformer stage, you are reducing the efficiency by using more stages because a transformer does the job of both copy stage as well as a filter stage
Tuesday, September 11, 2012
Clear the Job log in MetaData Repository
Sometime
having too much logs stored in the metadata repository can be a reason behind a
jobs slow performance.
Periodic log clearance would not only improve the performance but would keep the metadata database healthy.
Periodic log clearance would not only improve the performance but would keep the metadata database healthy.
Labels:
Administration
,
DataStage
,
delete
,
logs
,
Metadata
,
remove
,
Troubleshoot
,
Unix
,
windows
Subscribe to:
Posts
(
Atom
)