> ************ iinnttrroodduucciinngg XXSSLL ************ ********** ppoowweerrffuull ttrraannssffoorrmmiinngg ********** bbyy TTiimmoo ""BBllaazzkkoo"" BBooeewwiinngg,, _m_a_i_l_@_n_e_v_e_p_r_i_s_e_._d_e VVeerrssiioonn::0.01, 2002-05-03 >XML already has proven its strength, but besides using XML as data exchange format etcetera, you sometimes have to translate one XML format into another one. That might sound boring, but this article shall show, that transforming can be an interesting thing for every-day use. > ******** TTaabbllee ooff ccoonntteennttss ******** * _1_._0_ _o_n_ _t_r_a_n_s_f_o_r_m_i_n_g_ _X_M_L_ _-_ _w_h_y_ _a_t_ _a_l_l_? * _2_._0_ _i_s_n_'_t_ _t_h_i_s_ _j_u_s_t_ _f_o_r_ _m_i_d_d_l_e_w_a_r_e_ _/_ _i_n_t_e_r_f_a_c_e_ _d_e_s_i_g_n_e_r_s_? o _2_._1_ _c_r_o_s_s_ _m_e_t_h_o_d o _2_._2_ _i_n_s_i_d_e_ _o_u_t_ _m_e_t_h_o_d * _3_._0_ _a_b_o_u_t_ _X_S_L_ _a_ _b_i_t_ _m_o_r_e_ _i_n_ _d_e_t_a_i_l o _3_._1_ _r_e_q_u_i_r_e_d_ _s_o_f_t_w_a_r_e o _3_._2_ _p_r_a_c_t_i_c_a_l_ _e_x_a_m_p_l_e_s # _3_._2_._1_ _N_e_v_e_p_r_i_s_e_D_o_c_ _X_M_L # _3_._2_._2_ _C_o_n_t_e_n_t_ _M_a_n_a_g_e_m_e_n_t_ _S_y_s_t_e_m > ************ 11..00 oonn ttrraannssffoorrmmiinngg XXMMLL -- wwhhyy aatt aallll?? ************ _>_ _t_o_p_ _o_f_ _p_a_g_e As XML is already widenly used, it is no secret that XML is just a metalanguage but not an implementation. Thus, there co-exist several XML formats for one and the same purpose. For example: documentation. There are several formats and applications that use XML as storage language: * DocBook XML * XHTML * OpenOffice / Star Office 6.0+ * KOffice / KWord * MS Office (younger versions, since 2000?) * ... and NevepriseDoc XML ...just to name a few. However, just cos these formats use XML that does not mean that these formats can be opened by each application that is XML-aware. But as you know, XML laid the cornerstone for simple automated processing: you always have a structured tree, the data is well-formed.> And yes, there is an own XML implementation called XSL that is designed to convert one XML format into another one. XSL stands for eXensible Sylesheet Language. > ************ 22..00 iissnn''tt tthhiiss jjuusstt ffoorr mmiiddddlleewwaarree // iinntteerrffaaccee ddeessiiggnneerrss?? _>_ _t_o_p_ _o_f_ _p_a_g_e ************ Er, the second chapter is also the second question in this article, well, it was never desinged to be an FAQ... :-)> But to answer the question: yes and no. That depends on if you want to * interprete between more formats interchangeably (i call it the "cross method") * internally use one format but serve many (like to call it "inside out method") ********** 22..11 ccrroossss mmeetthhoodd ********** _>_ _t_o_p_ _o_f_ _p_a_g_e For what i am calling the cross method, just imagine a tool that was made to interprete beween several formats. A perfect utilization would be for example a converter in an office suite: it could import alien XML formats into the "own" one but also it could export the native format to the alien ones. ********** 22..22 iinnssiiddee oouutt mmeetthhoodd ********** _>_ _t_o_p_ _o_f_ _p_a_g_e The inside out method is quite similar, but is more one-way. It is about a centralized homogeneous system where you keep your data in just one format but want to serve many clients that each have their own format. You are right, it is just a cut-off version of the previous cross thing.> But it has a somewhat different nature and taste. The cross method will mostly be used in existing surroundings while the latter one is more for a new dedicated environment that wants as much bindings to the outer rims as possible. For an example, just see chapter four. > ************ 33..00 aabboouutt XXSSLL aa bbiitt mmoorree iinn ddeettaaiill ************ _>_ _t_o_p_ _o_f_ _p_a_g_e XSL is an XML implementation - that means XSL itself is a format that follows the XML notation guidelines (e.g. syntax, but also implements other XML derivates such as XLink and XPath). XSL's purpose is simple: to transform one XML document into another one. The name part "Stylesheet" might imply something different, but XSL is a real programming language - bound to one special task (transformation) - and features: * iteration (loops) * alternation (if and case) * variables and parameters * modularity (may include other XSL sources at runtime) During XML document translation, there are three instances involved: * an input document (source) * an output document (destination) * an XSL template file So when you want to transform the source document into a resulting destination format, an XSL template contains the rules on how (and if) information is shaped.> You can think of XSL as some kind of SSI (server side include) as known from nowadays web servers and CGI stuff: there you have an HTML template file that shall be filled with dynamically created content (but also statical file, useful for adding a persistent menu for example), e.g. a request in an search engine or web shop. Delivered data of a programm is bound into predefined places of the HTML file. Not so accurate examples of these technologies include: * Server Side Includes (e.g. in Apache) * server-side JavaScript * embperl (embedded Perl) * "HTML Business" (used in SAP software) XSL templates are very similar: the main part of an XSL template is written in the format of the destination document but filled with XSL statements for flow control and data selection. ********** 33..11 rreeqquuiirreedd ssooffttwwaarree ********** _>_ _t_o_p_ _o_f_ _p_a_g_e Of course you need special software to work with XSL, just as you need Perl to run Perl code.> This software is called an XSL processor and there are quite a few out there. On the one hand there are pure XSL processors, on the other hand there are quite many that have built-in processors: * Mozilla (built-in) * Internet Explorer (built-in) * Apache Cocoon (integral part of it) * Sablotron (pure but excellent XSL processor) As indicated in the previous chapter, this is the way these XSL processors work: (click image to enlarge) > > PPiicc 33aa:: creation process 1. reading and parsing an input document, 2. reading and parsing an XSL stylesheet, 3. applying rules (interpretation) and finally 4. creating a result tree (output document). ********** 33..22 pprraaccttiiccaall eexxaammpplleess ********** _>_ _t_o_p_ _o_f_ _p_a_g_e ******** 33..22..11 NNeevveepprriisseeDDoocc XXMMLL ******** _>_ _t_o_p_ _o_f_ _p_a_g_e We will not go into further technical details so far, but i want to show a very practical example for the "inside out" method.> What you are reading right now is the result of an XSL transformation (or XSL Transformation - XSLT). I have written this article in a custom XML format called NevepriseDoc XML. This is a format that i have developed myself over some months to perfectly fit my personal needs for this and other websites.> Why not HTML directly? There is one special reason: by writing in my XML format, i can create several output formats by using one source only, for example: * plain text * HTML w/ JavaScript * plain HTML * Postscript * PDF * ... You still might say: okay, that still works when writing in HTML; conversions to the named formats can still be archieved by it.> But there is one very important thing: HTML is a pure page description language. This means that i cannot describe internal parts a semantical way; in HTML i can say: make this text red and bold while in XML i can give a text a special meaning, e.g. this is a hint, this is a warning, this is a copyright note, this is an article's index, this is the description of an image and this is a list of definitions. In HTML i only can make these items by visual differences. But by using XML, i can also extract parts of the document, e.g. i just want to have all occuring source codes, just the definition lists or just the index. This is not possible in HTML except by manual editing. With NevepriseDoc XML, i write my articles once and can control the way this data is used. And without having redundant writings, i can easily create any output format. I just have to develop a XSL stylesheet once per format that are used to create the resulting documents.> Without the need to copy-and-paste an article, i quickly can do many more layouts and desings for the articles, some with JavaScript and some without. If over the time i want to do global design modifications, i just have to edit one single stylesheet and after automated recreation of the target documents, all of them carry the modification. This is very similar to Cascading Stylesheets (CSS) in HTML but with much more control. ******** 33..22..22 CCoonntteenntt MMaannaaggeemmeenntt SSyysstteemm ******** _>_ _t_o_p_ _o_f_ _p_a_g_e This would be the ideal way: an enterprise stores its important documents in an native XML format. Whenever you need a document, you can convert it to the desired format:> One needs it as a Word document, it can be converted into. One needs a web page for the internet www site, it can be converted. One needs a white paper, it can be converted.> And all that can be done by using one special resource - no redundant storage has to be done. The main thing is that you just not convert the document types but also gain much more control on the transformation and reserve sematic information in the XML source, while most target documents just contain layout information. > _>_ _t_o_p_ _o_f_ _p_a_g_e > _C_o_p_y_l_e_f_t (C)2002 by _T_i_m_o_ _"_B_l_a_z_k_o_"_ _B_o_e_w_i_n_g. This document is licensed under the terms of the _G_P_L. > >