A sneak peek to performing structural equation modeling using AMOS

Research often involves testing the relationship between variables in a hypothesis. While various quantitative techniques can be used for this purpose, it is structural equation modeling(SEM) approach that provides a visual & easy to interpret display of the causal relationship between the variables. 

Structural equation modeling, a combination of factor and multiple regression analysis, is a multivariate statistical analysis technique that evaluates the structural relationships between latent constructs and measured variables. Structural equation modeling is of two types.

  1. Measurement model 

This type of model represents the theory specifying how the measured variables group together to demonstrate the theory.

 

  • Structure model 

 

Here, the theory defining how the constructs are related to other constructs are represented. 

In research, SEM technique depends on the popular statistical software known as Analysis of Moment Structures (AMOS). This is because AMOS produces tabular outputs and graphic models using user-friendly tool. However, prior to performing SEM, it is a must to consider few assumptions such as:

  • Linearity – There should be a linear relationship between the endogenous and exogenous variables. 
  • Outlier – Since outlier impacts the model significance, the data should be free of outliers.
  • Multivariate normal distribution – Maximum likelihood approach is used for multivariate distribution. Also, the small changes in the multivariate results in the larger difference in chi-square test. 
  • Sequence – There exists a cause and effect relationship between the exogenous and endogenous variables.  
  • Model identification – The models must be exactly or over-identified as the under-identified models aren’t considered. 
  • Uncorrelated error terms – The error terms are uncorrelated with other variable error terms. 
  • Non-spurious relationship – This implies that the observed variance should be true. 

If all the above-mentioned assumptions hold good for SEM, AMOS then continues with the structuring equation modeling process by assuming that the data has been modeled. Some of the methods used by AMOS to perform SEM include:

  1. Generalized least squares – This method estimates the coefficients in the linear regression model if there exists correlation amongst the residuals. 
  2. Unweighted least squares – This approach estimates residual errors to access the conditional mean. 
  3. Browne’s asymptotic distribution free – This type of method is largely recommended when samples containing the non-normal data and SEM involves analysis of covariance structure. 

Model construction in AMOS

On determining the type of model, the next step is to run AMOS by clicking ‘start’ menu and choosing the ‘AMOS graphic’ option. When the AMOS starts running, a window known as ‘AMOS graphic’ appears, in which the user can manually draw the SEM model.

  • Data input – Choose the file name from the data file option and attach the data in AMOS for further SEM analysis. The user can also select this option by clicking the ‘select data’ icon. 
  • Observed variable – Use rectangle icon and draw the observed variables.
  • Unobserved variable – Deploy circle icon and draw the unobserved variables. 
  • Covariance – To denote the covariance between variables, select a double-headed arrow.
  • Cause-effect relationship – To establish the relationship between the observed and unobserved variable, use single-headed arrow in AMOS.  
  • Naming the variable – It is important to determine the variables to work with them precisely. Click on the variable in the graphical window, select ‘object properties’ option and name the variables in the AMOS. 
  • Error term – The error term appears next to the unobserved variable and is often used to draw the latent variable. 

After running the analysis, the outputs are displayed on the graphic window. However, the graphic window will display only the error term weights, standardized & unstandardized regressions. Some of the results produced by AMOS include:

  • Variable summary 

AMOS and the text output variable provides the option of viewing how many variables and which variables have been used for SEM analysis process. The user can also the number of observed and unobserved variables present in the model. 

  • Accessing the normality 

In the SEM model, data must be normally distributed. AMOS provides skewness, text outputs, Mahalanobis d-squared test, Kurtosis and also gives information about the normality of the data.

 Modification index 

The modification index describes the reliability of the path in the SEM model. If the modification index value is huge, then more paths can be added to the SEM model.

  • Estimates 

The estimate option in the AMOS text output will provide the output for regression weight, residual, standardised loading factor, covariance, indirect effect, direct effect, correlation, total effect, and many more. 

  • Error message 

If there is any error in the SEM model drawing process, then AMOS will either give an error message or will not calculate the result. 

  • Model fit 

The model fit will provide the result for goodness fit model statistics and present goodness fit indexes such as RMR, GFI, BCI, TLI, RMSER, and many more.

 Additionally, AMOS enables the functioning of SEM analysis and makes it easy to arrive at the statistics (where direct measurements are not possible). 

Chi-Square Test: Know about the Inferential Statistical test that Operates on Categorical Variables

Statistical test can be complex but it verifies & assures the quality of the study. In analytical work, the most crucial task is the comparison of data (sets of data) and making interpretations.

Inferential statistics, one among the two major branches of statistics, are concerned with making inferences based on the relationships found in the sample, to that in the population. This kind of statistics back the hypothesis, and such inferential test that supports the existence of group differences in the entire population is the Chi-Square test.

The Chi-Square test assumes (the null hypothesis) that the observed values match the expected values for categorical data. Chi-Square test is of two types,

  1. Chi-Square goodness-of-fit test – This test is performed for one categorical value and begins with hypothesizing that variable distribution behaves in a specific manner. For example, in order to identify the daily staffing needs of a store, the manager would wish to know if there is any consistency in the number of customers throughout the week.
  2. Chi-Square test for independence – This test is conducted for two categorical values. It compares variables in a contingency table & investigates if they are related. However, if the observed data doesn’t fit the model, the likelihood that the variables are dependent enhances, proving that the null hypothesis is incorrect.

For instance, this test can help us determine if the DRE is independent of biopsy results.

Chi-Square goodness-of-fit and test for independence depends only on the degrees of freedom and set of observed and expected values. These tests do not require any assumption pertaining to distribution of the parent population. However, if the test is performed along with the standard approximation, the applicability of Chi-Squared distribution, assumptions such as,

  • Simple random sample – It is a random sampling from a fixed population where every individual of the distribution/population of sample size has an equal probability of selection.
  • Constraints on cell frequency must be linear – If a sample with a large size is assumed, but the test is performed on smaller sample size, then the test will yield an inaccurate inference.
  • Theoretical cell frequency must be greater than 5 – As a rule of thumb, the expected cell count of a 2-by-2 table must be 5 or more and, for larger tables, it must be 5 or 80% of the cells.
  • Independence of observations – The observations must be independent of each other, which implies that Chi-Square cannot test correlated data such as matched pairs.

Chi-Square statistics 

Chi-Square test includes calculating a metric known as Chi-square statistic and the formula used here is,

Chi Square statistic

Where, E=(row total×column total) / sample size

Here the subscript ‘c’ indicates degrees of freedom, ‘O’ is the observed value and ‘E’ is the expected value. The summation indicates that the calculation must be done for each data item in the given data set.

Note that the Chi-square statistic cannot be used for percentages, proportions, and be used only on numbers.

For example, if you have 5% of 100 people, you must first convert 5% into the whole number and then run the test.

Chi-Square statistic includes variations which utilise the same concept, i.e., comparing observed & expected values. However, the type of variations used depends on how the data was collected and the hypothesis is being tested.

One of the most common forms of variation used for contingency tables is

Chi-Square statistic

Where E is the expected value, O is the observed value, and “i” is the ith position in the contingency table.

How to determine the statistical significance 

The low value for Chi-Square implies that the correlation between the two sets of data is high. Theoretically, if the observed & expected values are equal, then Chi-Square would be zero, which is practically impossible.

Determining if the test statistic is large enough to specify a statistically significant difference, includes two processes.

  1. Comparing calculated Chi-Square value and critical value from Chi-Square table. If the Chi-Square value is smaller than that of the critical value, there is no significant difference.
  2. Utilising P-value.

To calculate P-value manually,

  1. Firstly, determine the observed & expected values.
  2. This is followed by calculating the degrees of freedom using the formula,

Degrees of freedom = n-1, where ‘n’ is the number of variables.

  1. Compare observed & expected results with Chi-Square value (using the formula mentioned above).
  2. Choose a significance level (in decimals). Most commonly used value is P=0.05.
  3. Use Chi-Square distribution tables and approximate P value.
  4. Determine whether to reject the null hypothesis or not.

Example of Chi-Square test

Consider a random poll was conducted across 1000 voters (both male & female). The individuals who were classified based on their gender and whether they were supporters of party A, B or C.

Party A Party B Party C Total
Male 150 200 50 400
Female 270 300 30 600
Total 420 500 80 1000
  1. Calculate the expected value (E) using the above-mentioned formula
  2. In the second step, calculate the Chi-Square statistic
  3. Sum up all the values of Chi-Square statistics results, identify P-value and determine if the result is statistically significant or not

After obtaining the significant result, you can perform follow-up tests such as post-hoc test, to pinpoint the reasons for the significant result. To know more about it, visit https://www.rforge.net/doc/packages/NCStats/chisqPostHoc.html.

Wilcoxon Rank-Sum test : A Non Parametric Alternative to two Sample T-Test

Statistics, a scientific approach to analyzing numerical data, is employed to discover relationships among the phenomena to describe, predict and control their occurrence.

Statistics helps the researcher to acquire precise, steadfast and dependable findings. Although there are several statistical tests such as ANOVA, independent t-test, etc. to arrive at the right result, one must choose the test according to the type of study.

For instance, if one wants to investigate if the means of two or more groups are different from each other, then he/she must use the ANOVA test. On the other hand, if a researcher wants to test the relationship between categorical variables, the Chi-square test is to be used.

Similarly, for the comparison of means of two independent groups, the two-sample t-test is used. However, if the t-test doesn’t satisfy the requirements for two independent samples, then Wilcoxon Rank-Sum is used as it can offer the two independent samples drawn from populations with an ordinal distribution. This test does not assume known distributions, does not deal with parameters, and hence it is considered as a non-parametric test.

Wilcoxon Rank-Sum test also known as Mann-Whitney U test makes two important assumptions. That is the assumption of independence and equal variance. These assumptions are sufficient for determining if the two populations are different. Additionally, if we assume that the two populations are identical (except for a difference in location), then Wilcoxon Rank-Sum can be utilized as a test of equal means or medians.

Power calculation for Wilcoxon Rank-Sum test

Power is nothing but the probability of rejecting the null hypothesis when it is false. The power calculation for the Wilcoxon Rank-Sum or Mann-Whitney U test is similar to that of the two sample equal-variance t-test except a few modifications are made to the sample size based on the assumed data distribution.

The sample size ni| is equal to ni|= ni/𝑊,

where 𝑊 is known as the Wilcoxon adjustment factor, which is based on the assumed data distribution.

In general, the valid range for the probability of accepting a false null hypothesis is 0 to 1. However, different domains have different standards for setting power.

Sample size conditions 

While solving for sample size, the researcher must choose a condition that describes the constraints either on N1 or N2 or both.

  1. Equal (N1 = N2) – This condition is utilized when a researcher has equal sample sizes in each group. Since both sample sizes are solved at once, no additional sample size parameters are required here.
  2. Include N1, solve for N2 –  This condition is chosen to fix N1 at some value, and then solve only for N2. However, for some values of N1, N2 value that is large enough to acquire the desired power may be absent.
  3. Enter N2, solve for N1–  In case a researcher wants to fix N2 at some value, and then solve only for N1, this condition is used. In this case, too, N1 that is large enough to get the desired power might be absent for some values of N2.
  4. Enter R = N2/N1, solve for N1 & N2<span”> – To choose this condition, one must set a suitable value for the ratio of N2 to N1. This is followed by the determination of required N1 & N2 to obtain the desired power using PASS approach. An equivalent representation of R is
    N2 = R * N1.
  5. Include percentage in group 1, solve for N1 & N2 – Here, the researcher must set a definite value for the percentage of the total sample size in group1. Next, PASS determines the required N1 and N2 with the value of percentage entered to acquire the desired power.
  6. N1 (sample size, group 1) – This condition is used if group allocation = “Enter N1, solve for N2.” Where N1 is the number of individuals sampled from the group 1 population and must be equal or greater than 2. Here a single or a series of values can be entered.
  7. N2 (sample size, group 2) – If group allocation = “Enter N2, solve for N1,” this condition is utilized. Here N2 is the number of individuals sampled from the group 2 population and must be greater or equal to 2. A single or a series of values can be entered in this condition.

The Wilcoxon Rank-Sum test is less sensitive to outliers when compared to that of the two-sample t-test and valid for data from any distribution.

However, it reacts to other differences between the distributions such as differences in shape, especially if the focus is on the differences in location between the two distributions. This is considered as the major disadvantage of the Wilcoxon test. Also, when the assumptions of the two-sample t-test hold, this test is less likely to detect a location shift in comparison with the t-test.

Difference Between One Way ANOVA And Two Way ANOVA

When talking about research in the field of Science or Social science, whether it is Biology, Business, Economics, Psychology, Sociology, or any other subject, the Analysis of Variance (ANOVA) is an important statistical tool for analysing the data. The tool is used to compare and analyse the results of laboratories when more than one factor can be of influence and must be distinguished from random effects. Two folds of the technique lead the comparison; i.e., one way ANOVA and Two-way ANOVA.

ANOVA analysis the statistics on the basis of the hypothesis, either null or an alternate hypothesis. Since a hypothesis is an educated guess of the possible results of the cause-and-effect relationship, it will either result for the cause or against the purpose. The null hypothesis in ANOVA is valid when all the sample means don’t have a significant difference. Similarly, the alternate hypothesis is valid when at least one of the sample mean is different from the rest of the sample means.

As the names indicate of the two-fold techniques, the researcher takes only one factor in one way ANOVA and the researcher investigate two factors simultaneously in two way ANOVA. The former one is a hypothetical test, testing one-factor using variance whereas the later one is a statistical technique studying the influencing variables. However, the independent variables of both the types are proportional to the ways in their names.

One way ANOVA is based on the assumption of normal distribution of the sample population, the ratio level of the dependent variables, the independence of the samples, and the variance of the population. While two way ANOVA is also based on the assumption of normal distribution of the sample population but the measurement of the dependent variable is at a continuous level, unlike the variation in one way ANOVA. The two way ANOVA studies the inter-relationship between the influence of independent variables on dependent variables.

Two way ANOVA is often taken as an extended version of one way ANOVA as the former one has many advantages in the comparison of the later one.

Pilot Study: All Answers Are Here

These days most of the journals have the online submission system for manuscripts. This has certainly made it easier for the authors to send their papers for the review process. In all good journals, once you have submitted your paper, there is an online tracking system which allows authors to follow the progress of their paper in the review process.

After having submitted the manuscript, the authors go through a journey of stress and anxiety. This is obvious and because of this they keep checking the status of their manuscript. When they get confused or are not able to comprehend the status, they feel all the more perplexed. It also happens at times that the status update does not change for a long time. This also causes them to worry and get tensed. It is not a good idea to contact the journals/editors of the journals very often. If you have not received any notification or update from them for a fortnight, it is acceptable to ask for feedback through a formal mail to the editor. Often, journals take anything between a fortnight to sometimes even a month to get back with their first response to the author.

With some basic doubts and questions answered, authors would be able to have some control on their anxiety as they would know the meaning of a few things they did not know before.

Though there is a lot of subjectivity in each of the journal or publication house but some generalised situations and status may work for all and give you clarity to help you comprehend what different things mean when you verify the status of your manuscript:

1. Manuscript Submitted: This status means that the manuscript has been submitted successfully by the author and there is no other formality pending from the end of the author. From here on it isn’t sent for editing immediately. First the formatting is checked and verified before sending on for further processing.

2. Editor Invited: This status does not apply to all the journals. But those who follow this status, mean that an editor has been assigned for the same and his approval and acceptance is awaited.

3.With Editor: According to this status, the control of your paper has been handed over to an editor and his feedback is awaited on the same. If it clears this stage, it is further sent on for peer review. However, if it gets rejected by the editor here only, it means that it does not live up to the standards of the journal and should not be further forwarded for peer review.

4.Reviewer Invited: Like Step 2, this again is an optional step and not all journals may incorporate this step. After the feedback from the editor, if it is positive then, this status update means that the manuscript has been sent to reviewers and their acceptance is awaited.
There are more status we can know of which we will discuss in the next blog.